DevOps Engineer
Architect of Platform Engineering – AI Supercompute Infrastructure | Cloud & Engineering
Warsaw
Apply on the employer's siteRole description
Description & Requirements
Who we are looking for
🚀
Build the future of AI infrastructure with us!
We're a technology consulting firm designing and operating
next-generation AI supercomputing platforms
for some of the world's most ambitious organizations.
As our
Architect of Platform Engineering
, you'll own the full platform stack—from
bare metal and operating systems
to
cluster orchestration, job scheduling, and observability
—helping enterprise and public sector clients build AI infrastructure at scale.
🏆 As a
multiple-time NVIDIA Consulting Partner of the Year (EMEA)
, we offer
direct collaboration with NVIDIA engineering teams
, early access to the latest technologies, and the opportunity to work on projects that most engineers won't encounter for years.
💡This role combines
deep technical ownership
with
client-facing leadership
. You'll shape platform strategy, work closely with customer teams, and help define how large-scale AI infrastructure is built and operated.
💡 We're looking for a
hands-on Architect
who enjoys solving complex technical challenges and thrives in fast-paced, entrepreneurial environments.
What We Expect:
- 8+ years of hands-on infrastructure and platform engineering experience, including full ownership of production systems
- cluster architecture, control plane operations, custom controllers/operators, multi-tenancy, and large-scale fleet management
- Slurm experience or other HPC/AI workload scheduling: job queuing, fair-share scheduling, MPI integration
- Strong Linux internals knowledge: kernel tuning, cgroups, namespaces, NUMA topology, hugepages, and storage subsystems
- Familiarity with high-speed networking: InfiniBand, RoCE, RDMA; tuning for distributed training workloads
- Infrastructure as Code fluency: Terraform, Ansible, Helm or equivalent
- Demonstrated ability to lead technical engagements with enterprise clients. Translating ambiguous requirements into clear deliverables, managing stakeholders across seniority levels, and navigating complex organizational dynamics
- Entrepreneurial mindset. Comfortable operating with autonomy and moving fast without sacrificing rigor
- EU Work Permit
Bonus Experience
- Proven experience managing NVIDIA GPU infrastructure: driver lifecycle, CUDA toolchain, MIG/MPS partitioning, NVLink/NVSwitch topologies, and GPUDirect RDMA
- Familiarity with NVIDIA Base Command Platform, DGX SuperPOD, or CSP GPU cloud deployments
- Experience with DCGM, or other GPU profiling and telemetry tooling
- Prior consulting, professional services, or client delivery experience in an infrastructure or cloud practice
- Contributions to open-source platform tooling or CNCF ecosystem projects
Your future role
Area of responsibility:
- Operating System Layer. Own OS hardening, kernel tuning, driver management (NVIDIA CUDA, OFED, MIG), and node lifecycle automation across heterogeneous GPU fleets
- Monitoring & Observability. Build and evolve monitoring, alerting, and telemetry stacks (Prometheus, Grafana, DCGM Exporter, OpenTelemetry) to deliver deep visibility into cluster health, GPU utilization, and job performance
- Reliability Engineering. Define SLOs, drive postmortem culture, and lead incident response for production AI compute infrastructure; treat reliability as a systemic, architectural property, not just dashboards and runbooks
- Platform Strategy & Advisory. Translate complex technical requirements into platform roadmaps and architectural recommendations tailored to each client's business context and maturity level
- Cluster Orchestration. Design, deploy, and operate Kubernetes and Slurm clusters at scale across client environments; own the full lifecycle from provisioning to decommission, including upgrades, rollbacks, and capacity planning
You will define how some of the most demanding AI compute environments in Europe and beyond are built and operated across a portfolio of clients rather than a single environment. No two engagements are the same. The problems here, GPU fleet reliability at scale, distributed training fault tolerance, enterprise-grade platform governance are unsolved at the frontier, and you will have our firm's senior network behind you to tackle them. If you want broad impact, technical depth, and the autonomy to build a platform engineering practice rather than just execute within one, this is the role.
What we offer
- Flexible working hours,
- Permanent employment or contract,
- Medical and health insurance,
- Multisport and other lifestyle benefits,
- Language courses,
- Friendly coworkers & team spirit,
- Multiple geographies and clients,
- Work for well-known brands,
- Exposure to trailblazing business and technology projects,
- A place in the first line of a digital transformation,
- Everyday opportunities to influence how and where we do our business,
- A development path to fit your needs.
Selection process
We kindly ask you to upload your CV in English.
Shortlisted candidates will be contacted for the interviewing process.
If your CV would be interesting for us there will be a few steps:
- Quick HR call or meeting;
- One or two HR and technical interviews with our colleagues from the respective team;
- Final decision.
About Deloitte
About the team
Our Cloud Engineering teams design and deliver interesting cloud projects for clients in Poland and abroad in areas of cloud development, DevOps, integration, migration, data management, infrastructure and others. We help our clients to strategize, design and implement and migrate solutions with use of modern cloud technologies.
Engineering | Deloitte Global
This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.