Skip to content
Platform Engineer

Platform Engineer

Generative AI Pipeline Engineer (Tech Lead)

CapsLock

Remotefulltime

Apply on the employer's site

Role description

CapsLock is a global IT marketing company that pioneers unique, scalable customer acquisition solutions for B2C clients. We solve complex sales challenges across a wide range of industries, primarily for large North American partners. By integrating our expertise in Digital Marketing, IT, Design, AI-driven analytics, proprietary MarTech, and sales consulting, we deliver powerful, data-informed customer acquisition solutions.

At CapsLock, people, technology, and forward-thinking innovation lie at the heart of everything we do. Our diverse, global team, fluent in over 10 languages, thrives on open collaboration and the exchange of new ideas, pushing boundaries to ensure our clients' success.

Role Overview
We are building CapsLock's AI Media Production platform, and we're hiring the tech lead who will architect it.

We generate visuals of real products, not concepts.
Every asset must match the client's actual product 1:1 - down to hardware and geometry. No artifacts, no hallucinations, no "close enough"
. This is a control and reproducibility problem, not a prompting problem, and it needs an engineer.

This is a greenfield build. You'll be the first engineer in this direction, working directly with the Design Director as the technical interface between creative teams and software infrastructure.

What You Will Work On

  • Architect and run our AI Media Production platform: serving ComfyUI or similar engines, GPU environments, one-click UIs for non-technical teams;
  • Ship modular, versioned, reproducible generative pipelines, node-based and/or code-first (e.g. diffusers);
  • Develop one-click presets and use cases for image and video generation;
  • Train LoRAs end-to-end: from real product photography to evaluated, production-ready models;
  • Build control-first generation into every pipeline: ControlNet, IP-Adapter, inpainting, reference conditioning;
  • Engineer QC into the pipeline itself: automated checks against product references, before a human ever reviews;
  • Build our creative intelligence loop: encode what makes our best-performing assets work into reusable presets and styles;
  • Own the technical roadmap with the Design Director; qualify new models and tools; drive adoption through enablement, not mandates.

Requirements:

Your must haves
You don't need to meet every single requirement. If you're excellent in most of these areas and can show shipped work, we'd love to hear from you.

  • Shipped generative pipelines adopted and used daily by other people;
  • Production-grade workflows in ComfyUI or similar technologies, or equivalent code-first experience (e.g. diffusers);
  • Strong Python: automation, API integrations, custom nodes;
  • LoRA training end-to-end, from dataset curation to evaluation;
  • Control techniques across modalities (t2i, i2i, i2v, v2v): ControlNet, IP-Adapter, inpainting, reference conditioning;
  • Deploying pipelines as internal tools;
  • Enough visual literacy to set the quality bar (no design portfolio required);
  • Confident English (B2+).

Nice to have

  • 3D/render-hybrid workflows: Blender or CAD renders as geometric ground truth with generated environments;
  • LLM orchestration and agentic workflows: automated prompt generation, VLM-based QC;
  • Experience producing visuals for performance marketing.

What We Offer

  • Remote Work - we offer a truly and fully remote environment. You choose where you are the most productive and comfortable to have an impact.
  • Paid vacations - generous paid vacation policy to ensure you have time to recharge
  • Unlimited Sick Days - we understand that being only human means getting sick or feeling under the weather from time to time, so we guarantee you time off as long as you need to recover and get back on your feet.
  • Ongoing Learning - people at CapsLock are deeply inquisitive and eager to learn new knowledge and skills, that's why we support and create learning opportunities like the Free Books program, workshops, conferences, and more.
  • Home Office - we will cover the equipment and furniture expenses to make sure you have the best work-from-home experience.
  • Physical Well–Being - We will cover costs of your medical insurance. Additionally we offer an allowance for a flexible fitness program.
  • Fun Stuff - there is never a shortage of fun stuff

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Platform Engineer

Senior Database Reliability Engineer (DBRE) (worldwide remote)

CloudLinux

Remotefulltime

Apply on the employer's site

Role description

CloudLinux / TuxCare is a remote-first infrastructure and security company. More than 300 engineers build and operate products used by hosting providers, enterprises, and internal service teams worldwide. Our Infrastructure Department runs the platforms behind CloudLinux OS, Imunify, KernelCare, TuxCare ELS, and our engineering systems.

We are hiring a
Senior Database Reliability Engineer
to join the Infrastructure DBA cell. This is a hands-on production ownership role, not a narrow ticket-processing DBA position. You will keep critical database services reliable, automate repeated work, support engineering teams, and reduce single-person dependency in our PostgreSQL, ClickHouse, MongoDB, and Redis operations.

PostgreSQL is the main requirement. ClickHouse experience is a strong plus, but it is not a day-one blocker. We need a senior engineer with enough database, Linux, automation, and incident-response depth to learn our ClickHouse environment quickly and operate it safely.

Your Responsibilities:

  • Own production PostgreSQL reliability: HA design, Patroni, PgBouncer, replication, failover, upgrades, vacuum/bloat control, query tuning, locks, indexes, capacity, backups, PITR, and restore validation
  • Improve disaster recovery and operational evidence: tested restores, documented recovery paths, measurable RTO/RPO targets, runbooks, and safe maintenance plans
  • Support the wider database estate: ClickHouse, MongoDB, and Redis. You will troubleshoot incidents, review access and data-safety changes, improve monitoring, and learn the production ClickHouse patterns already in use
  • Automate DBA workflows with Ansible, Terraform/OpenTofu, GitLab CI/CD, scripts, and reproducible runbooks for provisioning, grants, backups, restores, health checks, and ownership metadata
  • Help build DBaaS-style self-service capabilities so engineering teams can request databases, access, credentials, and operational checks with less manual DBA intervention
  • Improve observability and incident response through Grafana, metrics, logs, SLOs, alert rules, Opsgenie routing, and clear communication during production issues

What Success Looks Like:

  • PostgreSQL clusters have tested backup and restore paths, useful dashboards, clear ownership, and documented failover procedures
  • Repeated DBA tickets become automation or self-service workflows
  • ClickHouse operational knowledge is no longer a single-person dependency
  • Database incidents have owners, runbooks, evidence, and measurable recovery paths
  • Product and engineering teams get database help faster without sacrificing safety, auditability, or reliability

Why CloudLinux?

  • You will work on real production infrastructure used across CloudLinux and TuxCare products.
  • You will have a direct impact on reliability, incident response, developer experience, and operational resilience.
  • You will also work in an AI-assisted engineering culture where automation, documentation, Claude, Codex, and careful human verification are part of the daily operating model

Requirements
What We Expect From You:

  • Deep hands-on PostgreSQL experience in business-critical production environments, typically 5+ years or equivalent depth
  • Strong understanding of PostgreSQL internals and operations: MVCC, WAL, transactions, locks, indexes, query planning, replication, autovacuum, bloat, major upgrades, backups, PITR, and restore testing
  • Proven experience with highly available databases and the ability to reason about quorum, split-brain risk, failover, rollback, and recovery
  • Strong Linux and infrastructure fundamentals: systemd, networking, storage, filesystems, CPU/memory/disk bottlenecks, TLS, DNS, firewalls, and root-cause troubleshooting
  • Automation skills with Ansible and scripting. Terraform/OpenTofu, GitLab CI/CD, and merge-request based delivery are strong advantages
  • Ability to support more than one database engine. You do not need to be a ClickHouse expert on day one, but you must be ready to learn it quickly and take responsibility for it
  • Practical use of AI engineering assistants such as Claude and Codex. We expect you to use them to improve speed and quality, while personally verifying generated SQL, commands, scripts, and operational conclusions
  • English - upper-intermediate or higher - to ensure clear communication of progress within the teams

Nice to Have:

  • ClickHouse operations: replication, Keeper/ZooKeeper, MergeTree engines, distributed DDL, grants, row policies, backups, query troubleshooting, and cluster recovery
  • MongoDB replica sets and Percona Backup for MongoDB
  • Redis/Sentinel and broker/cache failure modes
  • Database observability, SLOs, golden signals, alert tuning, and executable incident runbooks
  • Building internal platforms, self-service portals, or DBaaS workflows for engineering teams

Benefits
What's in it for you?

  • A focus on professional development
  • Interesting and challenging projects
  • Fully remote work with flexible working hours, which allows you to schedule your day and work from any location worldwide
  • Paid 24 days of vacation per year, 10 days of national holidays, and unlimited sick leaves
  • Compensation for private medical insurance
  • Co-working and gym/sports reimbursement
  • Budget for education
  • The opportunity to receive a reward for the most innovative idea that the company can patent

By applying for this position, you agree with
CloudLinux Privacy Policy
and give us your consent to maintain and process your personal data with this respect. Please read our Privacy Policy for more information.

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Platform Engineer

Torq HyperAutomation Platform Engineer

Black Birch Group

Remotefulltime

Apply on the employer's site

Role description

We are looking for an experienced Torq HyperAutomation Platform Engineer to own, operate, and continuously improve the hyperautomation platform supporting our security operations.

This role is ideal for a hands-on security automation professional who can design, build, test, and maintain workflows for alert enrichment, triage, containment, notifications, case management, and incident response across multiple client environments.

The Torq HyperAutomation Platform Engineer will work closely with SOC analysts, incident responders, security engineers, and client technical teams to identify repetitive manual processes and convert them into reliable, secure, and well-documented automations. This person will serve as the primary engineer responsible for the Torq platform, its integrations, platform health, and automation roadmap.

Responsibilities

- Serve as the primary owner of the Torq hyperautomation platform, including administration, workflow development, access management, integration management, versioning, troubleshooting, and platform health.

- Design, build, test, deploy, and maintain automated security workflows for alert enrichment, triage, containment, escalation, notification, case management, and incident response.

- Build and maintain integrations with security platforms, including CrowdStrike Falcon, email security tools, identity providers, SIEM platforms, cloud services, threat intelligence sources, and ticketing systems.

- Automate CrowdStrike Falcon activities involving detections, host information, Real Time Response, host containment, event streaming, and CrowdStrike Next-Gen SIEM.

- Integrate ticketing, case management, and IT service management platforms such as ConnectWise, Jira Service Management, or ServiceNow into automated security workflows.

- Automate ticket creation, enrichment, assignment, routing, escalation, status updates, and closure.

- Develop custom integrations using REST APIs, webhooks, JSON, and scripting when native connectors are unavailable or do not meet operational requirements.

- Implement and troubleshoot API authentication methods, including OAuth 2.0, API keys, bearer tokens, and service accounts.

- Automate email security operations such as phishing alert intake, message tracing, email search and purge actions, user-reported phishing workflows, threat enrichment, and analyst notifications.

- Partner with SOC analysts and incident responders to identify high-volume, repetitive, and low-value manual tasks that can be converted into automations.

- Support the onboarding of new clients, security tools, credentials, integrations, and data sources into the automation platform.

- Establish standards for workflow naming, documentation, reusable components, secrets management, logging, testing, error handling, deployment, and change control.

- Build automations with appropriate failure handling, retry logic, logging, monitoring, and analyst notifications.

- Maintain workflow documentation, technical procedures, integration references, runbooks, and knowledge base articles.

- Monitor workflow performance, platform availability, integration health, errors, and automation success rates.

- Measure and report on automation impact, including alerts automatically enriched or triaged, analyst hours saved, workflow success rates, and reductions in response times.

- Continuously review and improve existing workflows to increase reliability, security, maintainability, and operational value.

Requirements

- 3+ years of experience in security operations, security engineering, integration engineering, automation engineering, or a related technical role.

- Hands-on experience working with a SOAR or security hyperautomation platform such as Torq, Palo Alto Cortex XSOAR, Splunk SOAR, Swimlane, Tines, or a comparable platform.

- Strong understanding of SOC processes, including alert intake, enrichment, triage, escalation, containment, investigation, and incident response.

- Strong working knowledge of REST APIs, including authentication, HTTP methods, status codes, request and response handling, pagination, rate limits, and error handling.

- Experience testing and troubleshooting API integrations using Postman, curl, or similar tools.

- Experience integrating with CrowdStrike Falcon APIs or a comparable endpoint detection and response platform.

- Experience automating EDR activities such as detection enrichment, endpoint queries, host actions, containment, or response actions.

- Experience working with email security platforms such as Microsoft Defender for Office 365, Proofpoint, Abnormal Security, Avanan, Barracuda, IRONSCALES, or similar tools.

- Experience building or supporting automated phishing investigation and response workflows.

- Experience integrating ticketing, case management, PSA, or ITSM systems such as ConnectWise, Jira Service Management, or ServiceNow.

- Scripting proficiency in Python, JavaScript, or a similar programming language.

- Strong experience working with JSON, webhooks, structured data, conditional logic, and data transformation.

- Ability to troubleshoot workflow failures involving authentication, permissions, API behavior, data formatting, platform configuration, and automation logic.

- Solid understanding of secrets management, secure credential handling, role-based access control, and change management practices.

- Strong analytical, troubleshooting, and problem-solving skills.

- Strong written and verbal communication skills in English.

- Ability to communicate technical concepts clearly to SOC analysts, engineers, leadership, and client stakeholders.

- Highly organized, detail-oriented, and able to manage multiple workflows, integrations, and priorities independently.

- Strong documentation skills and the ability to create clear technical procedures, runbooks, and workflow documentation.

Preferred Qualifications
- Direct experience administering and building production workflows in Torq.

- Experience working in an MSSP, MDR provider, SOC, or multi-tenant security environment.

- Experience managing automations across multiple client environments, credentials, security platforms, and technology stacks.

- Familiarity with Microsoft 365, Microsoft Entra ID, and Microsoft Graph API for identity and email automation.

- Experience with SIEM platforms such as CrowdStrike Next-Gen SIEM, CrowdStrike LogScale, Microsoft Sentinel, Splunk, or similar platforms.

- Experience working with SIEM query languages and security event data.

- Familiarity with threat intelligence and enrichment platforms such as VirusTotal, AbuseIPDB, URLScan, AlienVault OTX, or similar services.

- Familiarity with Git, workflow versioning, testing, peer review, and controlled production deployment practices.

- Experience creating reusable automation components, templates, and standardized integration patterns.

- Experience measuring automation performance and presenting operational metrics to technical or business stakeholders.

- Relevant certifications such as CrowdStrike CCFA, CrowdStrike CCFR, CompTIA Security+, GIAC certifications, or comparable security certifications.

Benefits

Why Join Black Birch Technology Group

• Work on cutting edge AI and automation initiatives

• Build enterprise grade solutions with real business impact

• Collaborate with innovative teams and clients across industries

• Opportunity for growth within a rapidly evolving AI and automation practice

• Flexible hybrid and remote work environment

- Competitive salary package.

- Paid sick days.

- Continuous training and professional growth opportunities.

- Performance-based incentives.

- Private health insurance.

- Christmas bonus.

- Supportive culture that values employee well-being.

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Platform Engineer

Staff Software Engineer (Ruby or GOlang)

Workato

Remotefulltime

Apply on the employer's site

Role description

About Workato
Workato delivers enterprise infrastructure for the agentic era, redefining iPaaS and helping enterprises unify data, applications, processes, and AI into a single, governed platform. A leader in Enterprise MCP and trusted by 50% of the Fortune 500, Workato’s cloud-native architecture connects every application, data source, and process to power real-time orchestration at scale. With enterprise-grade security and continuous innovation at its core, Workato provides the trusted foundation for organizations to automate with confidence and operationalize AI across the business. To learn more, visit www.workato.com

Why join us?
Ultimately, Workato believes in fostering a
flexible, trust-oriented culture that empowers everyone to take full ownership of their roles
. We are driven by
innovation
and looking for
team players
who want to actively build our company.

But, we also believe in
balancing productivity with self-care
. That’s why we offer all of our employees a vibrant and dynamic work environment along with a multitude of benefits they can enjoy inside and outside of their work lives.

If this sounds right up your alley, please submit an application. We look forward to getting to know you!

Also, Feel Free To Check Out Why:

  • Business Insider named us an “enterprise startup to bet your career on”
  • Forbes’ Cloud 100 recognized us as one of the top 100 private cloud companies in the world
  • Deloitte Tech Fast 500 ranked us as the 17th fastest growing tech company in the Bay Area, and 96th in North America
  • Quartz ranked us the #1 best company for remote workers

Responsibilities
We are looking for an exceptional Staff Backend Developer (Ruby/Go) to join our growing Engine team. The Engine team develops and maintains most things related to Workato Recipe runtime. Everything related to recipe execution: DSL, pulling events, processing webhooks, executing jobs. There are various aspects to it: performance, scaling, storage, durability, atomicity, concurrency guarantees, data protection, and encryption.

In This Role, You Will Also Be Responsible To:

  • We also consider strong Ruby-only candidates, as well as Golang-only candidates who are open to learning new languages.
  • Be a key contributor to the architecture and design of distributed applications with strict requirements for performance and resilience.
  • Build/extend/troubleshoot/fix complex heterogeneous Golang/Ruby applications, as well as small self-contained Golang microservices.
  • Write well-designed, testable, efficient code in Ruby and Golang.
  • Integration of data storage solutions Postgres/S3/Redis/Kafka/ClickHouse etc.
  • Contribute in all phases of the development lifecycle.
  • Provide code reviews to your teammates.
  • Provide technical leadership. Work with other teams on shared projects.
  • Evaluate and propose improvements to existing systems.
  • Identify bottlenecks and bugs, and devise solutions to these problems.
  • Help maintain code quality, organization and automatization.

Requirements
Qualifications / Experience / Technical Skills

  • Strong experience in building scalable distributed backend applications (7+ years).
  • Excellent understanding of distributed systems patterns and algorithms.
  • Great understanding of all building blocks of large web applications: databases, load balancers, application servers, message brokers, caching, monitoring, etc.
  • Excellent understanding of network protocols and stacks.
  • Excellent understanding of DB technologies: classic databases and modern no-SQL.
  • Knowledge of all common basic data structures and algorithms and how they are used is a must.
  • Multilingual programming experience: our code base is primarily in Ruby, with a trend to migrate to GOlang and Rust. At least two languages are required.
  • At least basic understanding of cloud deployments (k8s, Terraform, ArgoCD)
  • Experience of working with public cloud infrastructure providers(AWS/Azure/Google Cloud).
  • Excellent debugging, analytical, problem solving, and social skills.
  • BS/MS degree in Computer Science, Engineering or a related subject, 7+ years of industry experience.

Soft Skills / Personal Characteristics

  • Background in GOlang, Rust, WASM, Kotlin-multiplatform.
  • Background in network programming.
  • Background in application, data security.
  • Deep knowledge of physical DB design.
  • Experience of working with Docker and other isolation technologies.
  • Experience in related fields (DevOps, ML, DBA, Enterprise applications, etc).
  • Experience in building/deploying data processing pipelines is a plus.
  • Experience of working with third-party REST APIs at scale (request throttling, batch processing etc).

Soft Skills / Personal Characteristics

  • Ability to technically lead projects. Work with requirements, cost analysis.
  • Readiness to work remotely with teams distributed across the world and timezones.

(REQ ID: 2582)

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Platform Engineer

DevOps / Senior Cloud Infrastructure Engineer - Freelancer

Monterail

Remotefulltime

Apply on the employer's site

Role description

Job Description
We are looking for a DevOps / Senior Cloud Infrastructure Engineer to join us on a freelance basis to build the observability and incident management backbone for autonomous kitchen operations.

4-5 month project | full time | remote
What We're Looking For

  • Senior-level experience as an AWS Cloud Engineer or Site Reliability Engineer.
  • Strong, demonstrable focus on metric-driven observability, monitoring, and alerting at scale.
  • Fluent, hands-on experience in Python for tooling and automation.
  • Hands-on experience with Terraform for Infrastructure as Code.
  • A proven track record of designing, architecting, and owning production systems end-to-end.
  • Ability to work completely independently without a detailed spec sheet or heavy direction.
  • Experience with Jira Service Management or similar ITSM/incident platforms is a plus.
  • Experience with Grafana dashboarding pipelines at scale is a plus.
  • Exposure to AI-assisted ops tooling (AI-Ops, runbook automation) is a plus.
  • Prior experience in a fast-moving hardware, robotics, or IoT fleet environment is a plus.

What you'll do

  • Standardize and own the Tech Ops Incident Management Platform using Jira Service Management across our production fleets.
  • Automate incident resolution workflows with AI tooling, including runbook generation and assignment.
  • Design and implement proper documentation for all Tech Ops incident processes to ensure a clean handover.
  • Build out fleet management and task automation to support global 24/7 remote operations.
  • Own and customize the data-driven observability layer via Grafana across all internal tech teams.
  • Work closely with key leadership stakeholders (Head of Infrastructure & Cloud, VP Engineering) to independently drive architecture decisions.

Requirements
What do we mean by freelance?
Read more at
Monterail Tech Network

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

650 more openings in this category and country

Generative AI Pipeline Engineer (Tech Lead)CapsLock · North Macedonia

Apply on the employer's site