Skip to content
Site Reliability Engineer

Site Reliability Engineer

Ingénieur·e Linux (X/F/M)

ORNESS

Montreuil

Apply on the employer's site

Role description

À propos de l’offre d’emploi
Localisation :
Montreuil

Type de contrat :
CDI

Rémunération :
65K€ – 75K€

Télétravail :
50%

Contexte
Vous interviendrez pour un acteur majeur du secteur bancaire sur l’administration et l’évolution d’une infrastructure de 50 000 serveurs Linux dont une partie utilisée pour le trading où les maîtres mots sont « la performance et la sécurité »

L’infrastructure est particulièrement diversifiée en environnement Linux avec aussi bien de la RedHat, de la Debian et de l’Ubuntu sous différentes versions.

L’automatisation est principalement réalisée avec du Ansible (il reste un peu de Puppet et de Satellite). Vous devrez maîtriser parfaitement Ansible afin d’en repousser ses limites du fait de la volumétrie et de l’hétérogénéité de l’infrastructure

Il s’agit d’un contexte OnPrem avec du serveur physique comme du VMWare au sein d’une équipe d’expert·e·s Linux soudé·e·s avec peu de turnover

Vos missions

  • Gestion de l’obsolescence de l’infrastructure avec le patch management et les mises à jour de sécurité sur une infrastructure de 50 000 serveurs. L’environnement est particulièrement volumineux et hétérogène pour un contexte OpenSource
  • Migration de plusieurs milliers de serveurs RedHat vers Debian, sans compromettre la performance et la sécurité
  • Mise en place de virtualisation avec Proxmox
  • Intégration de Kubernetes (OpenShift) avec les équipes CaaS en charge de le déployer
  • Automatisation de l’installation et des mises à jour de l’infrastructure avec Ansible (avec AAP idéalement)
  • Gestion des incidents système N3 avec le troubleshooting (perf, stockage, système, identification, …)
  • Travail dans un contexte bancaire hautement sécurisé avec des impératifs de performance et de sécurité

Technos & environnements

  • Linux & Systèmes : Linux (RHEL, Debian, Ubuntu)
  • Container & Orchestration : Kubernetes, OpenShift
  • Automatisation & CI/CD : Ansible, Puppet, Satellite, Terraform, Bash, Python
  • Virtualisation : Proxmox, VMware
  • HPC : HPC

Profil recherché
Nous recherchons un profil d’expert·e Linux passionné·e par le logiciel libre et capable de descendre dans les couches basses pour trouver quotidiennement des solutions et améliorer l’infrastructure, le tout dans un cadre bancaire

Si les challenges techniques vous stimulent et que vous aimez travailler dans un environnement où la liberté va de pair avec l’exigence, ce poste est fait pour vous.

Qui sommes-nous ?
ORNESS est une ESN spécialisée dans les infrastructures critiques, les systèmes haute disponibilité et les environnements conteneurisés à forte performance. Nous accompagnons des acteurs majeurs de la finance et de l’industrie sur des projets où stabilité, sécurité et performance sont essentielles.

Processus de recrutement
Nous privilégions un processus simple et rapide :

  • Premier échange rapide avec l’équipe RH (5 à 10 minutes)
  • Entretien RH approfondi (environ 1 heure)
  • Entretien technique (environ 1 heure)
  • Rencontre avec le client (30 minutes à 1 heure)

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

Ingénieur en Automatisation et Systèmes Robotisés (H/F) - Canada

EIC - Experience Internship Canada

Parisfulltime

Apply on the employer's site

Role description

EIC est le leader du placement de stagiaires internationaux au Canada. Notre mission est d’accompagner les étudiants internationaux afin de leur offrir une expérience de stage enrichissante en Amérique du Nord dans des conditions optimales !

Pourquoi choisir Experience Internship Canada ?
🌎

✅ Plus de 200 entreprises partenaires

✅ 17 secteurs d’activité : Finance, Marketing, Informatique & Technologie, Immobilier, Ressources humaines, ONG, Ingénierie, Data Science, Sports…

✅ Leader sur le marché des stages internationaux au Canada (Montréal, Toronto, Vancouver,...)

Ce que notre programme de stage offre aux étudiants :

  • Un stage professionnel dans ton domaine
  • Hébergement en centre-ville (salle de sport, piscine,...)
  • Assistance dans les démarches de visa
  • Soutien local 24h/24 et 7j/7
  • Coaching individuel et test de personnalité MBTI
  • Accueil à l'aéroport
  • Formation professionnelle en entreprise
  • Accès à un réseau international de contacts
  • Événements culturels et sociaux (afterwork, match de hockey,...)
  • Cours de langues (anglais ou français - optionnel)

Les avantages d'un stage à l'international :

  • Valoriser son CV avec une expérience internationale reconnue par les employeurs
  • Développer de nouvelles compétences professionnelles
  • Améliorer son anglais et travailler dans un environnement multiculturel
  • Se démarquer sur le marché de l’emploi grâce à une immersion professionnelle à l’étranger
  • Avoir des responsabilités concrètes
  • Bénéficier d'un accompagnement personnalisé favorisant leur développement personnel et professionnel
  • Elargir son réseau professionnel à l’échelle internationale
  • Découvrir la culture nord-américaine

Consulter notre site web pour plus d'informations : https://www.eic\-canada.ca/

Contexte (Stage au Canada)

Ce stage vous permettra de participer à des projets de conception et d’optimisation de systèmes robotisés, tout en vous familiarisant avec les technologies modernes de robotique. Vous aurez l’opportunité d’évoluer dans un environnement dynamique où vous travaillerez sur des applications concrètes et innovantes en robotique et automatisation.

🛠 Missions principales

  • Conception et programmation des systèmes robotisés
    : Participer à la conception et à la programmation des systèmes robotisés en utilisant des outils comme
    ROS
    (Robot Operating System),
    Python
    , ou
    C++
    pour l’automatisation de processus.
  • Optimisation des processus automatisés
    : Aider à l'optimisation des systèmes automatisés pour améliorer l'efficacité, la rapidité et la précision des robots dans des environnements industriels.
  • Tests et validation des robots
    : Contribuer à la réalisation de tests sur les robots, en analysant leurs performances et en ajustant les paramètres pour garantir leur bon fonctionnement dans des conditions réelles.
  • Gestion des capteurs et actionneurs
    : Participer à l'intégration et à la gestion des capteurs (caméras, LiDAR, capteurs de proximité) et des actionneurs (moteurs, servomoteurs) utilisés dans les systèmes robotisés.
  • Veille technologique et recherche d’innovations
    : Se tenir informé(e) des dernières tendances et technologies en automatisation et robotique pour proposer des solutions innovantes aux projets en cours.
  • Support à l’intégration des robots dans les chaînes de production
    : Aider à l’intégration des robots dans des processus de production automatisés, en collaborant avec les équipes techniques pour assurer une mise en œuvre efficace.

Compétences recherchées :

  • Automatisation et Robotique
    : Expérience avec des systèmes automatisés et logiciels de simulation comme
    MATLAB
    ,
    SolidWorks
    ,
    ANSYS
    .
  • Programmation et Contrôle
    : Langages
    C
    ,
    C++
    , et
    MATLAB
    pour le développement de systèmes robotiques et d’automatisation.
  • Multidisciplinarité
    : Capacité à travailler avec des ingénieurs de diverses spécialités pour créer des solutions globales et performantes.

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

Site Reliability Engineer, Mistral Cloud

Mistral

Parisfulltime

Apply on the employer's site

Role description

About Mistral
Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector—co-creating customized AI systems that they can run on their terms.

We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited.

The Role
As a
Site Reliability Engineer (SRE)
on the
Cloud Platform
team, you will shape the reliability, scalability, and performance of our Cloud platform and customer-facing applications. You’ll work closely with software engineers and product teams to ensure our systems meet and exceed the expectations of both internal and external customers.

This role is critical in maintaining the stability and efficiency of our infrastructure, enabling seamless experiences for users and developers. Your expertise will directly impact the robustness of our AI platform, ensuring it operates at scale with minimal downtime.

What You Will Do

  • Design, build, and maintain scalable, highly available, and fault-tolerant infrastructures to support our Cloud platform.
  • Operate systems and troubleshoot issues in production environments, including interrupts, on-call responses, and infrastructure scaling.
  • Implement and improve monitoring, alerting, and incident response systems to minimize downtime and optimize performance.
  • Develop and maintain workflows and tools for CI/CD, containerization, orchestration, monitoring, and logging.
  • Participate in on-call rotations to respond to incidents and perform root cause analysis to prevent recurrence.
  • Drive continuous improvement in infrastructure automation, deployment, and orchestration.
  • Collaborate with software engineers to enable safe and reproducible model-training experiments.
  • Build and enhance a cloud platform that abstracts infrastructure complexities for science and engineering teams.
  • Design and develop new workflows, tooling, and automation to improve system reliability, availability, and performance.
  • Ensure infrastructure adheres to security best practices and compliance requirements in collaboration with the security team.
  • Document processes and procedures to ensure consistency and knowledge sharing across the team.

What We're Looking For

  • A Master’s degree in Computer Science, Engineering, or a related field.
  • 5+ years of experience in a DevOps or SRE role, with strong expertise in bare metal infrastructure and distributed systems.
  • Hands-on experience with site reliability issues, including root cause analysis, in-production troubleshooting, and on-call rotations.
  • Proficiency in working with reliability KPIs, such as observability, alerting, and SLAs.
  • Experience with CI/CD, containerization, and orchestration tools like Docker and Kubernetes.
  • Knowledge of monitoring, logging, alerting, and observability tools such as Prometheus, Grafana, ELK Stack, or Datadog.
  • Familiarity with infrastructure-as-code tools like Terraform or CloudFormation.
  • Proficiency in scripting languages (Python, Go, Bash) and a strong understanding of software development best practices.
  • Solid grasp of networking, security, and system administration concepts.
  • Excellent problem-solving and communication skills, with the ability to work effectively in a collaborative environment.
  • Experience in an AI/ML environment, high-performance computing (HPC) systems, or modern AI-oriented solutions (e.g., Fluidstack, Coreweave, Vast) is a plus.

What we offer
We offer a comprehensive benefits package designed to support your well-being, growth, and work-life balance. Benefits vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks.

For the most up-to-date details on benefits available in your location, please refer to our Benefits page.

Privacy Policy
Your privacy matters to us. You can learn more about how we handle your personal data in our Applicant Privacy Policy.

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

Site Reliability Engineer

Mistral

Parisfulltime

Apply on the employer's site

Role description

About Mistral
Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector—co-creating customized AI systems that they can run on their terms.

We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited.

The Role
As a
Site Reliability Engineer (SRE)
on the
Platform
team, you will shape the reliability, scalability, and performance of our platform and customer-facing applications. You’ll work closely with software engineers and research teams to ensure our systems meet and exceed the expectations of both internal and external customers.

This role balances day-to-day operations on production systems with long-term software engineering improvements. Your work will reduce operational toil, foster reliability, and ensure high availability for our web services, inference environments, and ML workloads. You’ll enable seamless replication of work environments across multiple HPC clusters, directly impacting the stability and efficiency of our AI platform.

What You Will Do

  • Design, build, and maintain scalable, highly available, and fault-tolerant infrastructures to support web services and ML workloads.
  • Ensure our platform, inference, and model training environments are always highly available and enable seamless replication across HPC clusters.
  • Operate systems and troubleshoot issues in production, including interrupts, on-call responses, and infrastructure scaling.
  • Implement and improve monitoring, alerting, and incident response systems to minimize downtime and optimize performance.
  • Develop and maintain workflows and tools for CI/CD, containerization, orchestration, monitoring, and logging.
  • Participate in on-call rotations to respond to incidents and perform root cause analysis.
  • Drive continuous improvement in infrastructure automation, deployment, and orchestration using tools like Kubernetes, Flux, and Terraform.
  • Collaborate with AI/ML researchers to enable safe and reproducible model-training experiments.
  • Build a cloud-agnostic platform that abstracts infrastructure complexities for science and engineering teams.
  • Design and develop new workflows, tooling, and automation to improve system reliability, availability, and performance.
  • Work with the security team to ensure infrastructure adheres to best practices and compliance requirements.
  • Document processes and procedures to ensure consistency and knowledge sharing across the team.

What We're Looking For

  • A Master’s degree in Computer Science, Engineering, or a related field.
  • 7+ years of experience in a DevOps or SRE role, with strong expertise in cloud computing and distributed systems.
  • Hands-on experience with site reliability issues, including root cause analysis, in-production troubleshooting, and on-call rotations.
  • Proficiency in working with reliability KPIs, such as observability, alerting, and SLAs.
  • Experience with CI/CD, containerization, and orchestration tools like Docker and Kubernetes.
  • Knowledge of monitoring, logging, alerting, and observability tools such as Prometheus, Grafana, ELK Stack, or Datadog.
  • Familiarity with infrastructure-as-code tools like Terraform or CloudFormation.
  • Proficiency in scripting languages (Python, Go, Bash) and a strong understanding of software development best practices.
  • Solid grasp of networking, security, and system administration concepts.
  • Excellent problem-solving and communication skills, with the ability to work effectively in a collaborative environment.
  • Experience in an AI/ML environment, high-performance computing (HPC) systems, or modern AI-oriented solutions (e.g., Fluidstack, Coreweave, Vast) is a plus.

What we offer
We offer a comprehensive benefits package designed to support your well-being, growth, and work-life balance. Benefits vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks.

For the most up-to-date details on benefits available in your location, please refer to our Benefits page.

Privacy Policy
Your privacy matters to us. You can learn more about how we handle your personal data in our Applicant Privacy Policy.

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

Senior Site Reliability Engineer (Platform Engineering)

Criteo

Parisfulltime

Apply on the employer's site

Role description

What You'll Do:

Within the Platform department of R&D, the Product Reliability Engineering (PRE) group acts as the bridge between Product Engineering, Platform Engineering and Infrastructure. The Platform PRE group comprises eight teams helping R&D design, build, and operate large-scale distributed systems reliably and efficiently. The common objective of the PRE teams is to build the most reliable platform in AdTech.

As a Senior Site Reliability Engineer within the PRE-Platform team, you will help Platform teams improve the reliability, scalability, and operability of their services and CI/CD pipelines. You will develop and maintain tooling and libraries (Python, Jenkins, Chef), support observability and performance initiatives, and lead technical migrations across infrastructure and core dependencies. You will also participate in an on-call rotation to help ensure production stability.

What You’ll Do

  • Own and improve the full lifecycle of production services — from design and deployment to operation and continuous improvement.
  • Partner with development teams before launch through system design reviews, platform and framework development, capacity planning, and production readiness assessments.
  • Improve reliability, scalability, and performance by automating operations and driving infrastructure and platform enhancements.
  • Maintain and optimize live systems through monitoring, observability, performance analysis, and incident management.
  • Participate in incident response and contribute to a culture of blameless postmortems and continuous learning.
  • Stack: C#, Java, Scala, Python, Go, Prometheus, Grafana, Kibana, Linux, Kubernetes, Mesos, and more.

Who You Are:

  • Strong software engineering experience with at least one modern programming language.
  • 5+ years of experience in Site Reliability Engineering, Platform Engineering, Software Engineering, or DevOps roles.
  • Experience designing, operating, and troubleshooting large-scale distributed systems in production environments.
  • Solid understanding of systems engineering fundamentals, including compute, networking, storage, and observability.
  • Hands-on experience debugging production issues, optimizing system performance, and automating operational workflows.
  • Ability to write clean, maintainable, and reliable code used in production systems and internal platforms.
  • Strong analytical and problem-solving skills, with the ability to collaborate effectively across engineering teams.
  • Curiosity, ownership, and a pragmatic mindset toward improving reliability and developer productivity.

Take a look at our R&D blog on Medium for access and insight into our engineering culture and achievements.

We acknowledge that many candidates may not meet every single role requirement listed above. If your experience looks a little different from our requirements but you believe that you can still bring value to the role, we’d love to see your application!

Who We Are:

We’re Criteo, the Commerce Intelligence Platform. Criteo helps businesses turn shopper signals into commerce outcomes while delivering more relevant experiences for shoppers. We use proprietary commerce intelligence and AI decisioning to drive relevance for shoppers and performance for businesses.

At Criteo, our culture is as unique as it is diverse. From our offices across the globe or from the comfort of home, our 3,600 Criteos collaborate together to build an open, impactful, and forward-thinking environment.

We foster a workplace where everyone is valued, and employment decisions are based solely on skills, qualifications, and business needs—never on non-job-related factors or legally protected characteristics.

What We Offer:

🏢 Ways of working – Our hybrid model blends home with in-office experiences, making space for both.

📈 Grow with us – Learning, mentorship & career development programs.

💪 Your wellbeing matters – Health benefits, wellness perks & mental health support.

🤝 A team that cares – Diverse, inclusive, and globally connected.

💸 Fair pay & perks – Attractive salary, with performance-based rewards and family-friendly policies, plus the potential for equity depending on role and level.

Additional benefits may vary depending on the country where you work and the nature of your employment with Criteo.

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

79 more openings in this category and country

Ingénieur·e Linux (X/F/M)ORNESS · France

Apply on the employer's site