Skip to content
Site Reliability Engineer

Site Reliability Engineer

CDI - Site Reliability Engineer (H/F)

Hermès

Parisfulltime

Откликнуться на сайте работодателя

Описание вакансии

À propos :
Artisans contemporains depuis 1837, nous devenons également artisans du code !
Hermès Digital développe, maintient et met à disposition la plateforme e-commerce de la Maison Hermès dans 34 pays à travers le monde. Dans un contexte d’hyper-croissance, et face aux enjeux qui en découlent, nous avons souhaité développer une
nouvelle solution e-commerce full-custom
et
orientée micro-services
afin de répondre au mieux aux besoins de nos clients. Ce projet de grande envergure est une opportunité unique pour créer un système e-commerce
from scratch
.

Nous voulons faire ressentir au travers du code et de l’architecture de cette plateforme la qualité et l’excellence que reflètent les créations Hermès. Pour ce faire, le projet sera mené selon les meilleurs pratiques de conception :
TDD, DDD
,
architecture hexagonale
… Le tout en
pair programming
pour permettre à chacun d’apprendre, de transmettre et d’évoluer !

Descriptif du poste :
En tant que
SRE
, au sein de l’équipe
H
ermès
D
igital
I
nfrastructures
HDI
et sous la responsabilité du Responsable Infrastructure, vous jouerez un rôle essentiel dans la garantie de la fiabilité et des performances des services numériques de notre organisation.

Vous travaillerez en étroite collaboration avec l'équipe de développement logiciel pour concevoir, mettre en œuvre et entretenir des systèmes répondant à des normes élevées en matière de disponibilité et de performance.

Vos responsabilités comprendront la Build et l'optimisation de l'infrastructure, l’Observability des systèmes, l'automatisation des tâches, la gestion des incidents et la collaboration avec les équipes de développement afin d'améliorer la fiabilité des services au sein de notre organisation.

Vous devrez travailler en étroite collaboration avec l'équipe Platform Engineering afin de communiquer vos observations et contribuer à l'amélioration des processus et de l'utilisation des outils, en tirant parti de votre expérience dans la collaboration avec les équipes de développement logiciel.

Vos missions :
Un SRE sera intégré à une ou plusieurs équipes de développement (Features teams) et sera donc responsable des Backlog d'infrastructure relatifs à ses équipes.

Afin de garantir le bon déroulement de sa mission quotidienne, les SRE est responsable de :

  • Gérer les sujets liés à l’infrastructure dans les backlogs des équipes de fonctionnalités dont il est responsable.

Le SRE travaille en étroite collaboration avec le PO, le Technical Leader et l’équipe technique pour comprendre les objectifs de l’équipe et définir la stratégie technique locale pour l’infrastructure.

Le SRE participe aux différents rituels des équipes de fonctionnalités dont il est responsable.

  • Gérer la CAF (Capacité A Faire) consolidée pour le Backlog d'infrastructure de chaque équipe de fonctionnalités dont ils sont responsables

Le SRE est chargé de prendre en compte la capacité (CAF) avec tous les facteurs liés à ses Backlog en collaboration avec les Product Owners (PO) et les Technical Leaders de ses équipes. Cela implique de prendre en considération divers aspects tels que les exigences métier et techniques et d'autres facteurs pertinents pour prendre les bonnes décisions concernant la gestion des Backlogs.

  • Prioriser le traitement des différentes demandes de Build à travers les différents Backlogs

Le SRE définit – en collaboration avec les PO et les Technical Leaders – et communique l'ordre dans lequel les différentes demandes de Build sont exécutées pour ses différents Backlogs. Cette priorisation garantit que les tâches critiques sont traitées rapidement et dans le bon ordre, contribuant ainsi à la fiabilité globale et aux performances des différents services au sein de notre organisation.

  • Garantir la qualité de l'infrastructure déployée dans son domaine d'activité

Le SRE est responsable de la qualité de l'infrastructure qu'il déploie, y compris sa stabilité, sa sécurité et sa conformité aux meilleures pratiques et aux normes de notre organisation.

  • Construire l'infrastructure selon les besoins

Lorsqu'une nouvelle infrastructure est nécessaire, le SRE est responsable de sa création et de sa configuration. Cela peut impliquer la configuration de serveurs, de bases de données, de réseaux ou d'autres composants selon les besoins.

  • Déléguer les tâches de Build en fonction des compétences et la maturité technique de ses équipes

Le SRE évalue l'expertise technique de ses équipes et attribue les tâches de Build en conséquence, veillant à ce que les membres de l'équipe travaillent sur des tâches conformes à leurs compétences et à leur expérience.

Le SRE s’engage dans une démarche de partage des connaissances, par le biais de sessions de peer programming ou autres.

  • Valider les Merge Request d'Infrastructure as Code (IaC) émanant des Devs

Le SRE
examine
et
approuve
les modifications apportées au code d'Infrastructure as Code (IaC). Cette étape de validation garantit que les modifications apportées à l'infrastructure sont cohérentes avec le code défini et n'introduisent pas de problèmes.

  • Développer et maintenir des systèmes de surveillance et d'alerte

Le SRE crée et gère des systèmes qui surveillent en permanence la santé et les performances des différents services et de l'infrastructure. Il configure également des alertes pour notifier les équipes en cas de problèmes potentiels ou d'incidents.

  • Collaborer avec les équipes de développement pour améliorer la fiabilité du système

Le SRE travaille en étroite collaboration avec les équipes de développement et l’équipe de Platform Engineering pour identifier et mettre en œuvre des améliorations visant à renforcer la fiabilité, la disponibilité et les performances des applications.

  • Planification et optimisation de la capacité

Le SRE évalue la capacité actuelle de l'infrastructure et planifient les besoins futurs. Ils optimisent l'allocation des ressources et la scalabilité de l’infrastructure pour garantir une utilisation efficace des ressources et une performance optimale des services.

  • Création de solutions d'automatisation pour les tâches répétitives

Afin de réduire le travail manuel et d'améliorer l'efficacité, le SRE automatise les tâches courantes, telles que la provision de serveurs ou la gestion des changements de configuration.

  • Fournir un soutien et une formation à l'équipe de développement

Le SRE aide les équipes de développement en offrant un soutien et une formation dans les domaines liés à l'infrastructure, à la fiabilité et aux meilleures pratiques.

  • Gérer et répondre efficacement aux incidents

Le SRE est responsable de la gestion et de la réponse aux incidents, veillant à ce que les problèmes soient résolus rapidement pour minimiser les temps d'arrêt et les interruptions.

  • Identifier et atténuer les goulets d'étranglement du système et les problèmes de performances

Le SRE identifie proactivement les goulets d'étranglement et les problèmes de performances au sein du système et prend des mesures pour les résoudre afin de maintenir des performances optimales du système.

Le SRE, en collaborant avec les Développeurs, contribue activement aux tests de performance pour identifier et résoudre proactivement les goulets d'étranglement potentiels et les problèmes de performances.

  • Établissement et promotion des meilleures pratiques en ingénierie de la fiabilité

Le SRE promeut et met en œuvre les meilleures pratiques dans le domaine de l'ingénierie de la fiabilité, encourageant une culture d'amélioration continue.

  • Conformité en matière de sécurité

Le SRE Veille à ce que l'infrastructure respecte les normes de conformité en matière de sécurité.

  • Planification du Disaster Recovery Plan

Développer et maintenir des plans de reprise après sinistre pour minimiser les temps d'arrêt en cas de défaillance du système.

  • Optimisation des coûts

Le SRE est responsable de la surveillance et de l’optimisation des coûts de l'infrastructure, y compris l'allocation et l'efficacité d'utilisation des ressources.

  • Documentation

Le SRE est responsable de la rédaction et de la mise à jour de la documentation relative à l'infrastructure, aux processus et aux meilleures pratiques pour faciliter le partage des connaissances et l'intégration des nouveaux membres de l'équipe.

Environnement technique :

  • Langages de programmation : PHP 8, Javascript
  • Framework : Symfony 5, NodeJs, ReactJs
  • Web services : RESTful
  • Cloud: AWS, Alibaba Cloud
  • Orchestration et conteneurs : Kubernetes, Docker
  • Automatisation : Terraform, Helm, Kostumize
  • Gestion des configurations : Ansible
  • Architecture événementielle : SQS, SNS, Kafka
  • Moteur de recherche : ElasticSearch
  • Bases de données : Postgresql, MySQL, MongoDB
  • Cache : Elasticache Redis / Memcache
  • Observabilité: Prometheus, Thanos, Loki, Tempo, Grafana
  • Artifactory: JFrog
  • CI/CD : Jenkins, Gitlab, ArgoCD
  • Security : HashiCorp Vault, OKTA

Bénéfices pour vous :
Vous rejoignez la Maison Hermès, artisan de produits d’exception !

Vous êtes au cœur d’un projet
from scratch
passionnant

Vous intégrez une équipe bienveillante soucieuse de la qualité de son code et de l’évolution de ses membres,

Vous bénéficiez d’une grande autonomie et vos prises d’initiatives sont encouragées.

Profil recherché :
Compétences Techniques :

  • Vous avez au minimum 3 ans d’expérience professionnelle en tant que SRE. Vous êtes adepte des méthodes agiles, méthodologie SRE et GitOps.
  • Vous avez une maîtrise approfondie de la plateforme AWS, Docker et Kubernetes (EKS).
  • L’Infrastructure as Code (Terraform, Ansible, Helm et Kustomize),
  • L’observabilité (Prometheus, Thanos, Loki, Tempo, Grafana)
  • Le CI/CD (Gitlab-CI, Jenkins, Sonarqube, ArgoCD),
  • La création d’environnements et la sécurité n’ont pas de secret pour vous.
  • Vous avez déjà mis en place et maintenu des services communs, tels que

Authorization Server (OpenID provider)

Event Bus/Messaging.

Vault

Vous pratiquez couramment l’Anglais (à l’écrit et à l’oral).

Compétences Comportementales :

Vous avez un
excellent sens relationnel
et vous êtes bon
communicant
. Vous avez une bonne
capacité d’adaptation,
le
souci du résultat
, le
sens du service
et l’
esprit d’équipe
. Vous êtes
curieux
et
rigoureux
. Enfin, vous avez l’envie et la capacité d’
auto-apprentissage,
vous cherchez à vous
améliorer en continu !
Employeur responsable, nous nous engageons dans l’éthique, les diversités et l’inclusion. Rejoignez l’aventure humaine Hermès !

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

Site Reliability Engineer

Site Reliability Engineer (m/f/d)

Allianz Partners

Saint-Ouenfulltime

Откликнуться на сайте работодателя

Описание вакансии

Key Responsibilities
As a Site Reliability Engineer within Advanced Analytics (DA3) in the Chief Data & AI Office at Allianz Partners, you will join the platform engineering team to own the reliability and operational health of the central engineering platform.

You will define and maintain service level objectives, drive incident response at the infrastructure layer, and systematically eliminate operational toil through automation.

You will work closely with Platform Engineers, Security Engineers, and incident-response leads to ensure the platform meets its reliability commitments across production workloads spanning AI services, Java APIs, and frontend applications.

Through this role, you will have the main following responsibilities:

  • Define, instrument, and maintain SLOs and SLIs for platform components; own error budget tracking and produce regular reliability reports for senior leadership.
  • Serve on the on-call rotation as the infrastructure escalation tier; lead incident response for cluster-level, network-level, and storage failures; chair blameless post-incident reviews.
  • Implement and operate Kubernetes infrastructure (AKS): cluster lifecycle management, networking, resource quotas, autoscaling configuration, and multi-tenancy patterns across product team namespaces.
  • Develop Infrastructure as Code (Terraform) to provision and manage Azure resources with consistency, auditability, and repeatable rollback capability.
  • Build and maintain observability infrastructure: Prometheus, Grafana, Azure Monitor, and Application Insights; own alerting rules, dashboards, and distributed tracing coverage across platform components.
  • Perform capacity planning and cost-aware resource management: right-size node pools, tune vertical and horizontal pod autoscalers, and identify resource waste across namespaces.
  • Identify and eliminate toil: automate repetitive operational tasks through scripting and tooling; measure and track toil reduction over time.
  • Maintain platform reliability procedures: rolling upgrades, backup and recovery testing, disaster recovery runbooks, and change freeze coordination.
  • Contribute to CI/CD pipelines and GitOps tooling (GitHub Actions, ArgoCD) from a reliability and deployment safety perspective; work with platform engineering on release gates and rollback mechanisms.
  • Collaborate with incident-response leads on incident SLA targets and operational procedures; work with Security Engineers on infrastructure hardening and vulnerability remediation.

What You Bring

  • 5+ years professional experience in site reliability engineering, DevOps, or platform engineering roles.
  • Strong Kubernetes experience: cluster operations, networking (Ingress, network policies), storage, autoscaling, and hands-on troubleshooting across production environments.
  • Solid Infrastructure as Code experience with Terraform; familiarity with Bicep or ARM templates is a plus.
  • Production experience with Azure cloud services: AKS, ACR, Key Vault, Azure Monitor, Application Insights, Virtual Networks, and Private Endpoints.
  • Strong observability experience: Prometheus, Grafana, centralized logging, alerting configuration, and distributed tracing instrumentation.
  • Working knowledge of SLO/SLI methodology: error budget principles, reliability target setting, and capacity planning.
  • Structured incident management experience: on-call ownership, blameless post-incident review, and runbook authorship.
  • Scripting and automation proficiency in Python or bash for toil elimination and operational tooling.
  • Strong CI/CD experience: GitHub Actions and ArgoCD or equivalent GitOps tooling.

Ways of Working

  • Comfortable in agile, iterative delivery environments with personal ownership and accountability for platform reliability.
  • Clear communicator across global, cross-functional stakeholders; able to translate technical reliability metrics into business impact for non-technical audiences.
  • Proactive learner with pragmatic adoption of AI-assisted developer tools (e.g., GitHub Copilot, Claude Code) to improve automation coverage and delivery velocity.

Nice to Have

  • Kubernetes certifications: CKA or CKAD.
  • Experience supporting AI or ML infrastructure workloads: GPU scheduling, model serving platforms, or inference pipeline operations.
  • Exposure to chaos engineering practices and fault injection testing.
  • FinOps experience: reserved capacity planning, resource right-sizing programs, and cost attribution per team or workload.
  • Service mesh experience (Istio, Linkerd) for traffic management and reliability patterns.
  • Experience in regulated industries (insurance, finance, healthcare) where auditability, change traceability, and secure-by-default operations are standard practice.

How We Hire
Allianz Partners does not accept unsolicited CV’s or approaches from agencies. We only work with partners on our approved supplier lists, under contract. Any unsolicited submission will not be considered.

What We Offer
Our employees play an integral part in our success as a business. We appreciate that each of our employees are unique and have unique needs, ambitions and we enjoy being a part of their journey. We are there to empower and encourage you with your personal and professional development ensuring that you take control by offering a large variety of courses and targeted development programs.

All that in a global environment where international mobility and career progression are encouraged. Caring for your health and wellbeing is key priority for us. This is why we build Work Well programs to providing you with peace of mind and give the flexibility in planning and arranging for a better work-life balance.

90377 | Data & AI | Professional | Allianz Partners | Full-Time | Permanent

Allianz Group is one of the most trusted insurance and asset management companies in the world. Caring for our employees, their ambitions, dreams and challenges, is what makes us a unique employer. Together we can build an environment where everyone feels empowered and has the confidence to explore, to grow and to shape a better future for our customers and the world around us.

At Allianz, we stand for unity: we believe that a united world is a more prosperous world, and we are dedicated to consistently advocating for equal opportunities for all. And the foundation for this is our inclusive workplace, where people and performance both matter, and nurtures a culture grounded in integrity, fairness, inclusion and trust.

We therefore welcome applications regardless of ethnicity or cultural background, age, gender, nationality, religion, social class, disability or sexual orientation, or any other characteristics protected under applicable local laws and regulations.

Join us. Let's care for tomorrow.

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

Site Reliability Engineer

Ingénieur en Automatisation et Systèmes Robotisés (H/F) - Canada

EIC - Experience Internship Canada

Parisfulltime

Откликнуться на сайте работодателя

Описание вакансии

EIC est le leader du placement de stagiaires internationaux au Canada. Notre mission est d’accompagner les étudiants internationaux afin de leur offrir une expérience de stage enrichissante en Amérique du Nord dans des conditions optimales !

Pourquoi choisir Experience Internship Canada ?
🌎

✅ Plus de 200 entreprises partenaires

✅ 17 secteurs d’activité : Finance, Marketing, Informatique & Technologie, Immobilier, Ressources humaines, ONG, Ingénierie, Data Science, Sports…

✅ Leader sur le marché des stages internationaux au Canada (Montréal, Toronto, Vancouver,...)

Ce que notre programme de stage offre aux étudiants :

  • Un stage professionnel dans ton domaine
  • Hébergement en centre-ville (salle de sport, piscine,...)
  • Assistance dans les démarches de visa
  • Soutien local 24h/24 et 7j/7
  • Coaching individuel et test de personnalité MBTI
  • Accueil à l'aéroport
  • Formation professionnelle en entreprise
  • Accès à un réseau international de contacts
  • Événements culturels et sociaux (afterwork, match de hockey,...)
  • Cours de langues (anglais ou français - optionnel)

Les avantages d'un stage à l'international :

  • Valoriser son CV avec une expérience internationale reconnue par les employeurs
  • Développer de nouvelles compétences professionnelles
  • Améliorer son anglais et travailler dans un environnement multiculturel
  • Se démarquer sur le marché de l’emploi grâce à une immersion professionnelle à l’étranger
  • Avoir des responsabilités concrètes
  • Bénéficier d'un accompagnement personnalisé favorisant leur développement personnel et professionnel
  • Elargir son réseau professionnel à l’échelle internationale
  • Découvrir la culture nord-américaine

Consulter notre site web pour plus d'informations : https://www.eic\-canada.ca/

Contexte (Stage au Canada)

Ce stage vous permettra de participer à des projets de conception et d’optimisation de systèmes robotisés, tout en vous familiarisant avec les technologies modernes de robotique. Vous aurez l’opportunité d’évoluer dans un environnement dynamique où vous travaillerez sur des applications concrètes et innovantes en robotique et automatisation.

🛠 Missions principales

  • Conception et programmation des systèmes robotisés
    : Participer à la conception et à la programmation des systèmes robotisés en utilisant des outils comme
    ROS
    (Robot Operating System),
    Python
    , ou
    C++
    pour l’automatisation de processus.
  • Optimisation des processus automatisés
    : Aider à l'optimisation des systèmes automatisés pour améliorer l'efficacité, la rapidité et la précision des robots dans des environnements industriels.
  • Tests et validation des robots
    : Contribuer à la réalisation de tests sur les robots, en analysant leurs performances et en ajustant les paramètres pour garantir leur bon fonctionnement dans des conditions réelles.
  • Gestion des capteurs et actionneurs
    : Participer à l'intégration et à la gestion des capteurs (caméras, LiDAR, capteurs de proximité) et des actionneurs (moteurs, servomoteurs) utilisés dans les systèmes robotisés.
  • Veille technologique et recherche d’innovations
    : Se tenir informé(e) des dernières tendances et technologies en automatisation et robotique pour proposer des solutions innovantes aux projets en cours.
  • Support à l’intégration des robots dans les chaînes de production
    : Aider à l’intégration des robots dans des processus de production automatisés, en collaborant avec les équipes techniques pour assurer une mise en œuvre efficace.

Compétences recherchées :

  • Automatisation et Robotique
    : Expérience avec des systèmes automatisés et logiciels de simulation comme
    MATLAB
    ,
    SolidWorks
    ,
    ANSYS
    .
  • Programmation et Contrôle
    : Langages
    C
    ,
    C++
    , et
    MATLAB
    pour le développement de systèmes robotiques et d’automatisation.
  • Multidisciplinarité
    : Capacité à travailler avec des ingénieurs de diverses spécialités pour créer des solutions globales et performantes.

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

Site Reliability Engineer

Site Reliability Engineer, Mistral Cloud

Mistral

Parisfulltime

Откликнуться на сайте работодателя

Описание вакансии

About Mistral
Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector—co-creating customized AI systems that they can run on their terms.

We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited.

The Role
As a
Site Reliability Engineer (SRE)
on the
Cloud Platform
team, you will shape the reliability, scalability, and performance of our Cloud platform and customer-facing applications. You’ll work closely with software engineers and product teams to ensure our systems meet and exceed the expectations of both internal and external customers.

This role is critical in maintaining the stability and efficiency of our infrastructure, enabling seamless experiences for users and developers. Your expertise will directly impact the robustness of our AI platform, ensuring it operates at scale with minimal downtime.

What You Will Do

  • Design, build, and maintain scalable, highly available, and fault-tolerant infrastructures to support our Cloud platform.
  • Operate systems and troubleshoot issues in production environments, including interrupts, on-call responses, and infrastructure scaling.
  • Implement and improve monitoring, alerting, and incident response systems to minimize downtime and optimize performance.
  • Develop and maintain workflows and tools for CI/CD, containerization, orchestration, monitoring, and logging.
  • Participate in on-call rotations to respond to incidents and perform root cause analysis to prevent recurrence.
  • Drive continuous improvement in infrastructure automation, deployment, and orchestration.
  • Collaborate with software engineers to enable safe and reproducible model-training experiments.
  • Build and enhance a cloud platform that abstracts infrastructure complexities for science and engineering teams.
  • Design and develop new workflows, tooling, and automation to improve system reliability, availability, and performance.
  • Ensure infrastructure adheres to security best practices and compliance requirements in collaboration with the security team.
  • Document processes and procedures to ensure consistency and knowledge sharing across the team.

What We're Looking For

  • A Master’s degree in Computer Science, Engineering, or a related field.
  • 5+ years of experience in a DevOps or SRE role, with strong expertise in bare metal infrastructure and distributed systems.
  • Hands-on experience with site reliability issues, including root cause analysis, in-production troubleshooting, and on-call rotations.
  • Proficiency in working with reliability KPIs, such as observability, alerting, and SLAs.
  • Experience with CI/CD, containerization, and orchestration tools like Docker and Kubernetes.
  • Knowledge of monitoring, logging, alerting, and observability tools such as Prometheus, Grafana, ELK Stack, or Datadog.
  • Familiarity with infrastructure-as-code tools like Terraform or CloudFormation.
  • Proficiency in scripting languages (Python, Go, Bash) and a strong understanding of software development best practices.
  • Solid grasp of networking, security, and system administration concepts.
  • Excellent problem-solving and communication skills, with the ability to work effectively in a collaborative environment.
  • Experience in an AI/ML environment, high-performance computing (HPC) systems, or modern AI-oriented solutions (e.g., Fluidstack, Coreweave, Vast) is a plus.

What we offer
We offer a comprehensive benefits package designed to support your well-being, growth, and work-life balance. Benefits vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks.

For the most up-to-date details on benefits available in your location, please refer to our Benefits page.

Privacy Policy
Your privacy matters to us. You can learn more about how we handle your personal data in our Applicant Privacy Policy.

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

Site Reliability Engineer

Site Reliability Engineer

Mistral

Parisfulltime

Откликнуться на сайте работодателя

Описание вакансии

About Mistral
Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector—co-creating customized AI systems that they can run on their terms.

We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited.

The Role
As a
Site Reliability Engineer (SRE)
on the
Platform
team, you will shape the reliability, scalability, and performance of our platform and customer-facing applications. You’ll work closely with software engineers and research teams to ensure our systems meet and exceed the expectations of both internal and external customers.

This role balances day-to-day operations on production systems with long-term software engineering improvements. Your work will reduce operational toil, foster reliability, and ensure high availability for our web services, inference environments, and ML workloads. You’ll enable seamless replication of work environments across multiple HPC clusters, directly impacting the stability and efficiency of our AI platform.

What You Will Do

  • Design, build, and maintain scalable, highly available, and fault-tolerant infrastructures to support web services and ML workloads.
  • Ensure our platform, inference, and model training environments are always highly available and enable seamless replication across HPC clusters.
  • Operate systems and troubleshoot issues in production, including interrupts, on-call responses, and infrastructure scaling.
  • Implement and improve monitoring, alerting, and incident response systems to minimize downtime and optimize performance.
  • Develop and maintain workflows and tools for CI/CD, containerization, orchestration, monitoring, and logging.
  • Participate in on-call rotations to respond to incidents and perform root cause analysis.
  • Drive continuous improvement in infrastructure automation, deployment, and orchestration using tools like Kubernetes, Flux, and Terraform.
  • Collaborate with AI/ML researchers to enable safe and reproducible model-training experiments.
  • Build a cloud-agnostic platform that abstracts infrastructure complexities for science and engineering teams.
  • Design and develop new workflows, tooling, and automation to improve system reliability, availability, and performance.
  • Work with the security team to ensure infrastructure adheres to best practices and compliance requirements.
  • Document processes and procedures to ensure consistency and knowledge sharing across the team.

What We're Looking For

  • A Master’s degree in Computer Science, Engineering, or a related field.
  • 7+ years of experience in a DevOps or SRE role, with strong expertise in cloud computing and distributed systems.
  • Hands-on experience with site reliability issues, including root cause analysis, in-production troubleshooting, and on-call rotations.
  • Proficiency in working with reliability KPIs, such as observability, alerting, and SLAs.
  • Experience with CI/CD, containerization, and orchestration tools like Docker and Kubernetes.
  • Knowledge of monitoring, logging, alerting, and observability tools such as Prometheus, Grafana, ELK Stack, or Datadog.
  • Familiarity with infrastructure-as-code tools like Terraform or CloudFormation.
  • Proficiency in scripting languages (Python, Go, Bash) and a strong understanding of software development best practices.
  • Solid grasp of networking, security, and system administration concepts.
  • Excellent problem-solving and communication skills, with the ability to work effectively in a collaborative environment.
  • Experience in an AI/ML environment, high-performance computing (HPC) systems, or modern AI-oriented solutions (e.g., Fluidstack, Coreweave, Vast) is a plus.

What we offer
We offer a comprehensive benefits package designed to support your well-being, growth, and work-life balance. Benefits vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks.

For the most up-to-date details on benefits available in your location, please refer to our Benefits page.

Privacy Policy
Your privacy matters to us. You can learn more about how we handle your personal data in our Applicant Privacy Policy.

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

Ещё 80 вакансий по этой категории в этой стране

CDI - Site Reliability Engineer (H/F)Hermès · France

Откликнуться на сайте работодателя