DevOps Engineer
Senior DevOps Engineer
Remote
Apply on the employer's siteRole description
We are seeking a
Senior DevOps Engineer
to join an AI Workbench Platform team focused on operationalizing domain foundation models. These are large models trained on log, seismic, drilling, and production data, analogous to general-purpose large language models but specialized for oil & gas subsurface and production domains. While these models exist today and are maturing, there is currently no platform to commercialize them, enable internal teams (geo units, business units, data scientists) to use them at scale, or allow external customers to interactively consume or fine-tune them. In this role, you will help design and build the infrastructure foundation that brings these powerful domain models to production.
This is a fully remote position that offers you the flexibility to work from any location in Armenia, whether it's your home or well-equipped offices in Yerevan or Gyumri.
Responsibilities
- Design, build, and maintain scalable infrastructure to support the AI Workbench Platform and its underlying domain foundation models
- Implement and manage Kubernetes clusters with multi-GPU scheduling capabilities to support large-scale model training and inference workloads
- Develop and maintain Infrastructure as Code using Terraform to provision cloud resources reliably and reproducibly
- Package, deploy, and manage applications using Helm and Kustomize across multiple environments
- Collaborate with data scientists, ML engineers, and business units to enable seamless model consumption and fine-tuning workflows at scale
- Ensure platform reliability, scalability, and security for both internal teams and external customers
- Optimize resource utilization and cost efficiency across GPU-intensive workloads
- Establish CI/CD pipelines and automation to accelerate platform delivery and model deployment
- Monitor system performance and troubleshoot production issues to maintain high availability
- Contribute to platform architecture decisions and best practices for MLOps at enterprise scale
Requirements
- 3+ years of experience in DevOps, Site Reliability Engineering, or Infrastructure Engineering roles
- Expertise in Kubernetes with proven experience in multi-GPU scheduling for AI/ML workloads
- Proficiency in Terraform for Infrastructure as Code and cloud resource management
- Skills in Helm and Kustomize for Kubernetes application packaging and configuration management
- Background in building and operating production-grade platforms that support large-scale, distributed workloads
- Understanding of MLOps principles and infrastructure requirements for training and serving large foundation models
- Capability to collaborate cross-functionally with data scientists, ML engineers, and business stakeholders
- Excellent command of written and spoken English (B2+ level)
Nice to have
- Prior experience with LightOps infrastructure
- Familiarity with on-premises infrastructure environments
- Knowledge of High-Performance Computing (HPC) systems and workloads
We offer
- We connect like-minded people
- Delivering innovative solutions to industry leaders, making a global impact
- Enjoyable working environment, whether it is the vibrant office or the comfort of your home
- Opportunity to work abroad for up to two months per year
- Relocation opportunities within our offices in 55+ countries
- Corporate and social events
- We invest in your growth
- Leadership development, career advising, soft skills and well-being programs
- Certifications, including GCP, Azure and AWS
- Unlimited access to EPAM's internal learning database
- Free English classes with certified teachers
- We cover it all
- Participation in the Employee Stock Purchase Plan
- Monetary bonuses for engaging in the referral program
- Comprehensive medical & family care package
- Four trust days per year for personal needs
- Discounts for fitness clubs
- Benefits package (hotels, restaurants, stores and services)
EPAM is global leader in AI transformation engineering and integrated consulting, serving Forbes Global 2000 companies and ambitious startups. With over thirty years of expertise in custom software, product and platform engineering, we empower our clients to become AI-Native enterprises, driving measurable value from innovation and digital investments.
This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.