Skip to content
DevOps Engineer

Список недели

Все такие вакансии — одним письмом

Сейчас вы читаете одну вакансию. Таких же на витрине сотни, и каждую неделю выходят новые. Выберите, что присылать, оставьте почту — список придёт сам, без поисков и без возвращения сюда.

Считаем, сколько вышло за прошлую неделю…

Первое письмо приходит сразу, дальше — раз в неделю. Отписка в один клик из любого письма, адрес больше никуда не уходит.

DevOps Engineer

Systems Engineer - m/f/d

Langdock

Berlin

Откликнуться на сайте работодателя

Описание вакансии

Where Europe's enterprises adopt AI
Langdock is the AI platform used by more than 10,000 companies to give employees secure access to the leading AI models, to build and share agents and to automate repetitive workflows. We have grown past $40M ARR while remaining a small team, and we care deeply about operating efficiently across the entire company.

For many enterprises, Langdock is becoming the place where most of the net-new work is produced. As people and agents create more documents, analyses, decisions, and automations inside AI interfaces, the context and data behind that work accumulate within Langdock. This gives us the opportunity to earn a larger role in their technology stack by building a platform they choose to rely on.

Our ambition is to build that platform for European enterprises while preserving their control over data, model providers, and deployment environments. We have made meaningful progress at the application layer, but much of the foundation beneath it still needs to be built.

You can watch the Meet the engineering team video to get a feeling for how we work.

The role
Systems Engineers solve our most complex technical problems inside and across Langdock's shared services when existing building blocks or standard architectures are insufficient. These are problems where correctness under failure, performance, scale, security, and cost interact.

This is a project-oriented role rather than ownership of one permanent layer. You might spend a period on distributed execution, then move into storage architecture, workload isolation, or model serving as company priorities change. The constant is sustained investigation of unfamiliar systems and responsibility for turning that understanding into a production system.

Systems Engineers own the difficult mechanism and the measurable improvement it creates, whether in correctness, capability, performance, reliability, or cost. They operate what they build, while the Platform owner remains accountable for the service contract and long-term lifecycle around it.

What You Might Work On
The systems agenda includes:

  • Distributed state and execution. Design systems that remain correct when work is concurrent, long-running, retried, resumed, or moved between processes. Define explicit invariants for ordering, idempotency, recovery, and tenant isolation.
  • Storage and data systems. Improve how large volumes of customer and agent-generated data are stored, versioned, moved, and recovered while preserving consistency, data residency, and predictable performance.
  • Secure compute for agents. Develop the isolation, runtime, and scheduling mechanisms behind a shared sandbox service for untrusted code. The system needs persistent filesystems, deny-by-default networking, scoped mounts, fast startup, suspension, resumption, and predictable scheduling without weakening the security boundary.
  • Inference systems. Improve the serving systems beneath the Model Gateway, optimizing throughput, latency, accelerator utilization, reliability, and cost. The work is measurement driven: understand the workload, identify the actual bottleneck, and decide where changes to serving, scheduling, caching, model format, or hardware create structural advantage.
  • Resource scheduling and isolation. Make CPU, memory, storage, network, and accelerator capacity explicit so one workload cannot degrade another. Improve placement, admission control, backpressure, autoscaling, and recovery across deployment environments.
  • Billing and usage metering. Build the concurrent metering mechanism behind the billing capability, accounting for heterogeneous units such as model tokens and sandbox compute time. It must stay correct when many workloads report concurrently, events arrive late or more than once, and long-running jobs reserve capacity before their final usage is known. This involves idempotent ingestion, atomic reservations, reconciliation, and auditable records.

You will start with one focused problem based on your experience and the team's priorities. The expectation is not broad activity across every domain; it is a substantial improvement in the capability, performance, reliability, or economics of the system you take on.

Tech stack

  • Go and TypeScript in one Bazel monorepo, with implementation choices driven by the system
  • Linux, containers, microVMs, filesystems, and networking
  • Protobuf and gRPC for service contracts
  • Kubernetes across GCP, AWS, Azure, and on-premises deployments
  • Terraform and Terragrunt for infrastructure orchestration
  • PostgreSQL and Redis where durable metadata or coordination requires them
  • Open-source model serving and accelerator infrastructure

You do not need prior experience with every item. You do need enough systems depth to enter an unfamiliar part of the stack, understand its behavior, and make consequential changes safely.

How We Work

  • We operate with high trust and autonomy in squads of 3 to 4 engineers. A squad owns its roadmap, prioritization, technical decisions, and operation in production. Engineers are expected to find the context they need, ask for input when it improves the outcome, and move work forward without waiting for every next step to be assigned.
  • We align asynchronously before scheduling a meeting. Product requirement documents (PRDs) define the user problem, intended outcome, and constraints. Design documents make architectural boundaries, tradeoffs, failure modes, migrations, and rollouts explicit. People read and challenge the thinking asynchronously; once the context is shared, a short in-office discussion or whiteboard session usually resolves the remaining questions quickly.
  • We optimize for leverage. Engineers choose the AI tools that work for them, supported by clear ticket context, focused branches, automated tests, and AI review before human review. We also invest in observability, migration tooling, automated recovery, and runbooks so recurring product maintenance does not depend on someone remembering a manual step.
  • The engineer who ships a change owns it in production. If something breaks, you lead the fix.

What We Are Looking For

  • You have owned technically difficult production systems in areas such as distributed systems, databases, runtimes, inference, networking, or storage. You operate what you build and remain responsible for it after deployment.
  • You reason from first principles, form testable hypotheses, and use measurement to understand unfamiliar behavior. You investigate beyond the first working solution until you understand the underlying mechanism.
  • Given an underspecified problem, you identify the most important invariant, decide where depth will change the outcome, and defend what you deliberately leave out.
  • You reason precisely about concurrency, isolation, ordering, data loss, and failure, while finding pragmatic ways to improve performance, reliability, or cost without unnecessary complexity.
  • You use AI tools as leverage while verifying their output and retaining ownership of the result. You communicate complex systems clearly, expose uncertainty, and work constructively with others.

Working here
This is an in-office role at Greifswalder Strasse 212 in Berlin. We work together in person because it helps us build trust, develop shared context, and make decisions quickly.

Most engineers start around 8:30. We usually eat lunch together, and dinner is available for people who stay later. Running and gym are part of the routine for many of us.

You need an existing right to work in Germany. We do not currently sponsor visas.

Compensation
The salary range for this role is €90,000–€140,000 gross per year, depending on level and scope. All roles include equity. Salaries are tied to levels, not negotiation.

We will figure out the right level together based on your experience and scope. Levels are about the work you own, not your title or years of experience.

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

Скоро на этой странице

Резюме под эту вакансию — и билет в розыгрыш

Мы разбираем объявление до настоящих требований и переписываем ваше резюме под него — вопросами, а не выдумкой: ни одна строка не появится без вашего подтверждения. Войдите, чтобы получить это первым, — и попасть в розыгрыш.

  • Резюме под конкретную вакансию, а не «универсальное»
  • Ответы хранятся: правится любой, а не весь разговор заново
  • Всё в аккаунте — открывается с любого устройства

Разыгрываем

Скидка на сопровождение

Победителей выбираем случайно среди заявок с подтверждённой почтой. Дата розыгрыша и полные правила — на странице розыгрыша.

Правила розыгрыша

DevOps Engineer

Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure

NVIDIA

Berlinfulltimesenior

Откликнуться на сайте работодателя

Описание вакансии

NVIDIA's Deep Learning Frameworks (DLFW) Infrastructure team is looking for a deeply technical Senior HPC Cluster Administrator to lead the design, deployment, and reliability of our large-scale GPU compute clusters. These systems run the most demanding deep learning training, inference, and high-performance computing workloads in the industry — from DGX/HGX platforms to ground-breaking Grace Blackwell systems. You will drive architectural decisions across compute, networking, and storage, and partner closely with software, research, and product teams to keep our infrastructure ahead of the workloads it supports.

What You'll Be Doing

  • Own the full lifecycle of GPU compute clusters — procurement, provisioning, configuration management, monitoring, and deprecation — across heterogeneous Linux environments (DGX, HGX, embedded systems)
  • Design and scale storage solutions (NFS, Lustre, WekaFS, or equivalent) with a clear roadmap for capacity and performance growth
  • Lead automation of infrastructure using modern IaC tools (Ansible, Terraform) and CI/CD pipelines (GitLab)
  • Manage and optimize job scheduling via Slurm, including fair-share policies, reservation management, and MIG/GPU partitioning strategies
  • Maintain and improve observability stacks (Prometheus, Grafana, DCGM) and drive proactive resolution of hardware and software incidents
  • Collaborate with ML engineers and software teams to tune cluster configuration for large-scale distributed training workloads
  • Evaluate and introduce new technologies — networking fabrics (InfiniBand, NVLink, EFA/RDMA), storage tiers, container runtimes — to improve performance and reliability
  • Mentor junior engineers and contribute to team-wide engineering standards

What We Need To See

  • BS/MS in CS, EE, CE, or equivalent hands-on experience
  • 5+ years of experience deploying and administering large-scale HPC or ML training clusters
  • Deep expertise in Linux systems administration at scale
  • Strong scripting and automation skills in Python and/or bash
  • Hands-on experience with Slurm (scheduling, accounting, cgroup configuration)
  • Proficiency with configuration management and IaC (Ansible required; Terraform a plus)
  • Experience with container technologies (Docker, Apptainer/Singularity, Kubernetes)
  • Solid understanding of high-speed networking (InfiniBand, RoCE, RDMA, EFA)
  • Experience with distributed/parallel filesystems and storage architecture
  • Ability to own problems end-to-end and communicate clearly with engineering and management stakeholders

Ways To Stand Out From The Crowd

  • Experience with NVIDIA GPU infrastructure tools (DCGM, nvidia-smi, MIG, NVSwitch diagnostics)
  • Familiarity with cluster management platforms (Colossus, Bright Cluster Manager, xCAT, or similar)
  • Experience supporting large-scale distributed deep learning workloads (PyTorch, JAX, Megatron)
  • Knowledge of BMC/IPMI/Redfish for out-of-band management and hardware lifecycle
  • Background in MLOps tooling or ML platform engineering

Join our team of world-class engineers and be part of the groundbreaking work we do at NVIDIA. We are committed to encouraging a collaborative and inclusive environment, where every team member has the opportunity to thrive and make a significant impact!

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. For Poland: The base salary range is 221,250 PLN - 383,500 PLN for Level 3, and 292,500 PLN - 507,000 PLN for Level 4. , , JR2015529

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

DevOps Engineer

DevOps Engineer *German required (GCP) (m/w/d)

ventx - we make IT!

Munichfulltime

Откликнуться на сайте работодателя

Описание вакансии

Zur Verstärkung unseres Teams im Großraum München suchen wir, zum nächstmöglichen Termin, engagierte Leute die Lust haben, deine bisher erworbenen IT-Kenntnisse zu erweitern und sich firmenintern zu einem DevOps Cloud Engineer (m/w/d) mit Schwerpunkt AWS Cloud ausbilden zu lassen.

Aufgaben

  • Lust sich in neue Themengebiete einzuarbeiten

  • Mut sich zum Spezialisten ausbilden zu lassen

  • CI/CD Implementationen

  • Erstellung von Infrastrucure as Code

  • Consulting von Projekten im Bereich IT-Infrastruktur

  • Gute Deutsch- und Englischkenntnisse

Qualifikation

Dein Profil

eine erfolgreich abgeschlossene IT-Ausbildung bzw. ein abgeschlossenes Studium der Informatik oder entsprechende Berufserfahrung

  • sicherer Umgang mit Linux

  • beherrschen mindestens einer Script-Sprache wie Bash oder Python

  • Kenntnisse in der Administration von Netzwerken

  • gute Infrastrukturkenntnisse und Coding-Skills

Wünschenswert, aber nicht zwingend erforderlich

Erfahrungen in dem Bereich von Cloud Providern – insbesondere GCP

  • Know-How über Infrastruktur als Code – Cloudformation, Terraform

  • Kenntnisse der Tools CI/CD, GIT, Ansible, Chef, Puppet, Jenkins

  • Kenntnisse in der Administration von Netzwerken

  • praktische Erfahrung mit Monitoring-Tools wie Grafana, Graphite, Prometheus sowie Logging-Tools wie ElasticSearch, lostack, Kibana

Benefits

  • neben attraktiver Vergütung und abwechslungsreicher Tätigkeit, Spaß bei der Arbeit

  • flache Hierarchien mit kurzen Entscheidungswegen

  • ein Team dass sich gerne bei Problemen und Fragen gegenseitig unterstützt

  • moderne Büroräume mit aktuellster Hard- und Software

  • Team-Events

  • Kostenlose Getränke im Office

BEWIRB DICH JETZT! Wir antworten in kürzester Zeit.

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

DevOps Engineer

Staff Infrastructure Engineer

World

Munichfulltimelead

Откликнуться на сайте работодателя

Описание вакансии

About The Company
Tools for Humanity (TFH) designs and builds technology behind World. World is building a real human network designed to accelerate people in the age of AI. As bots and autonomous agents reshape the internet, people, institutions, and applications need a trusted way to confirm who is a real human while preserving privacy. The TFH and World tech stacks make this possible: the Orb verifies real, unique people, World ID proves it privately, and World App puts these capabilities, and more, in people’s hands. Together, they add a human layer to an AI-driven internet.

World is already running at a global scale. More than 17 million people across 160 countries have verified with World ID, and more new Orb verifications take place each week. World App is already among the most used wallets globally. Developers are integrating World ID to build safer online experiences and create spaces where real people can participate, earn, and be recognized in ways AI simply can’t replicate.

Founded in 2019, TFH has more than 400 people across hardware, software, AI, cryptography, mobile engineering, and global operations. Our teams come from OpenAI, Tesla, SpaceX, Apple, Google, Stripe, Meta, Coinbase, Palantir and MIT Media Lab. We’re backed by leading investors, including a16z, Khosla Ventures, Bain Capital Crypto, Blockchain Capital, Variant, Tiger Global, and Coinbase Ventures, as well as prominent operators and founders across fintech and AI.

TFH and World have been featured on the cover of TIME Magazine, highlighted in Fast Company’s Next 5 in Fintech, and explored in a Bloomberg deep dive. The New York Times, Bankless and TechCrunch have all recognized our collective progress in identity, cryptography, AI, and global-scale hardware deployment. Our leadership is also named to the Time AI 100. Learn more about the newest product launches from our Liftoff event.

About The Team
Our Infrastructure team is a collaborative group of experienced engineers dedicated to supporting the World project's mission. With diverse backgrounds from across the technology landscape, we work together to solve complex challenges and provide a stable foundation for the entire organization. We are now thoughtfully expanding our team from Europe to the US, and we're looking for a new member to help us grow. We believe in sharing knowledge and supporting each other as we tackle the unique opportunity of building infrastructure for a global community.

This opportunity is based in Munich.
About The Opportunity
We are looking for a Staff Infrastructure Engineer to help establish our team's presence in the United States. This is a meaningful role where you will contribute significantly to the technical direction and best practices for our infrastructure in the region. Your work will involve a thoughtful blend of responsibilities, from hands-on systems administration to building reliable automation and contributing to our Site Reliability Engineering (SRE) practice. There is never a dull day; your contributions will support a wide variety of critical systems, including our anonymous multi-party computation platforms, internal security services, and blockchain node deployments. Ultimately, your work will directly support our teams and the project's long-term goals.

About You
With at least 10 years of experience, you are a senior infrastructure engineer who enjoys both teaching and learning from your peers.

  • You have a strong foundation in systems engineering, including hands-on experience with Linux, networking, and cloud environments like AWS or GCP.
  • You enjoy building tools and automation that make work easier and more reliable for everyone, with experience in infrastructure-as-code tools like Terraform.
  • You are a proactive and thoughtful problem-solver, taking the initiative to see challenges through to resolution.
  • You are a kind and clear communicator, able to work effectively with a distributed team and patiently support colleagues with different technical backgrounds.
  • You are motivated by the opportunity to contribute your expertise to a project with a broad and positive mission.
  • Your background includes experience working on large-scale, distributed systems.

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

DevOps Engineer

Senior Systems Engineer – Business Applications (m/w/d)

BWI GmbH

Strausbergfulltimesenior

Откликнуться на сайте работодателя

Описание вакансии

Sorge gemeinsam mit uns für die digitale Zukunftsfähigkeit der Bundeswehr.

Deine Aufgaben:

  • 2nd Level Support für die LAMP (Linux/Apache/MySQL/PHP) - Plattformen und Applikationen in unserer eigenen Cloud
  • Mitwirkung bei der automatisierten Bereitstellung von Plattformkomponenten durch CI/CD
  • Erstellung, Test und Implementierung von Containerbuilds
  • Bearbeitung von Störungen gemäß Spezifikation der BWI (Störungsannahme, Fehleranalyse, Fehlerbehebung)
  • Identifikation von Risiken, sowie Absicherung und Behebung von Security Risiken
  • Mitwirkung bei der Entwicklung von Schulungskonzepten
  • Pflege und Weiterentwicklung der technischen Dokumentation

Dein Profil:

  • Erfahrungen mit einem oder mehreren der folgenden Technologien: Webservern (Apache, NGINX), Linux Betriebssysteme (z.B. CentOS, SuSE, RHEL, Debian), Datenbanken (z.B.: MySQL, MariaDB, PostgreSQL), LAMP basierte Anwendungen
  • Gute Kenntnisse in Kubernetes und Cloud-nativen Konzepten
  • Kenntnisse von Anwendungsdeployments as Code via Helm-Charts, kustomize, k8s-Operators
  • Erfahrungen mit Automatisierungswerkzeugen (z.B. GitLab-CI, Terraform, Crossplane, Ansible, ArgoCD und FluxCD)
  • Erfahrungen mit Container-Builds (z.B. Docker, Buildah) und Artefaktverwaltung (Nexus, Harbor)
  • Vertrautheit mit Monitoring- , Logging- und Alerting-Stacks (z.B. Prometheus, Grafana, Instana)
  • Erfahrungen mit Skalierung und Ressourcenoptimierung in Container basierten Umgebungen sowie Git/Git-Ops
  • Grundverständnis TCP Netze wie Firewalling, Netzstruktur, Sicherheit
  • Erfahrungen mit CMDB/Asset-Managementsystemen
  • ITIL Grundkenntnisse (Change/Incident/Problem-Management)
  • Fließende Deutsch- und gute Englischkenntnisse

Wir bieten:

  • Durch abwechslungsreiche und gesellschaftlich relevante Aufgaben gewährleisten wir den reibungslosen IT-Betrieb und die Digitalisierung der Bundeswehr
  • Das Ziel eint uns. Dabei sind für uns ein wertschätzender Umgang miteinander sowie ein großer Teamgeist elementar
  • Die Vergütung liegt zwischen 57.800 € und 84.680 €. Die tatsächliche Höhe wird basierend auf deinem Verantwortungsbereich sowie Erfahrungen und Kompetenzen festgelegt
  • Wir bieten 30 Tage Jahresurlaub, 1 Brauchtumstag plus Optionen auf individuelle Anpassungen
  • Über unsere Benefit-App erhält man ein monatliches Guthaben und kann sich zusätzlich Steuervergünstigungen auf Tickets für den ÖPNV sichern
  • Wir ermöglichen Flexibilität, um Beruf und Privatleben in Einklang zu bringen, etwa durch mobiles Arbeiten oder Vertrauensarbeitszeit
  • Wir unterstützen die berufliche und persönliche Weiterbildung durch individuelle Maßnahmen sowie einen kostenfreien Zugriff auf LinkedIn Learning
  • Unser Jobradangebot ermöglicht das Leasing von bis zu 2 Fahrrädern

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

Ещё 237 вакансий по этой категории в этой стране

Systems Engineer - m/f/dLangdock · Germany

Откликнуться на сайте работодателя