Skip to content
Site Reliability Engineer

Site Reliability Engineer

Consultant Data Ops – F/H

Onepoint

Paris

Apply on the employer's site

Role description

Contribuez aux grandes transformations des entreprises et des acteurs publics en alliant innovation technologique et expertise métier, au service de nos clients et de la société pour les faire avancer durablement.

Au-delà de la RSE, nous avons développé notre propre approche, RESET, qui englobe l’ensemble de nos engagements en matière de responsabilité économique, sociale, environnementale et technologique, tant au sein de notre collectif que chez nos clients.

Onepoint, est une entreprise inclusive, handi accueillante, avec une politique de diversité affirmée.

Votre futur collectif :

  • Rejoignez notre communauté Data Intelligence, pour participer chez nos clients à la chaîne de valeur de la Data : Data Strategy, Data Management, Data Protection, Data Analytics, Data Platform, Data Science, IA Générative et bien d'autres...
  • Des missions variées dans tous secteurs confondus
  • Des bureaux inspirants au Trocadéro, avec vue sur la Tour Eiffel
  • Une vraie vie de communauté : Conférences, REX, Revolution Summit, Datapéros…

Vos nouveaux challenges pour ce poste :
Vous serez au cœur des pratiques DataOps, en automatisant les déploiements, en mettant en place la supervision et en garantissant la qualité des données.

Votre mission : appliquer les principes DevOps et SRE à l’univers data pour assurer des solutions robustes, résilientes et observables.

Vous intervenez entre autres sur les missions suivantes :

  • Automatiser les déploiements (CI/CD) des pipelines et plateformes data.
  • Mettre en place la supervision, l’observabilité et l’alerting data.
  • Appliquer les principes DevOps et SRE à l’écosystème data.
  • Gérer la qualité des données (tests, contrôles, SLAs).
  • Optimiser la fiabilité et la résilience des traitements.
  • Collaborer étroitement avec les équipes Data et Cloud.

Votre profil et les expertises que vous souhaitez développer :

  • Vous avez une expérience en DataOps, DevOps ou SRE appliqué à des environnements data.
  • Vous maîtrisez les concepts d’automatisation, monitoring et observabilité.
  • Vous souhaitez approfondir vos compétences sur les plateformes cloud (Azure, AWS, GCP) et les outils de data engineering.
  • Vous êtes reconnu(e) pour votre rigueur, votre sens de la qualité et votre capacité à fiabiliser des systèmes complexes.
  • Vous aimez travailler en équipe et contribuer à la mise en place de standards robustes pour l’écosystème data.

Technologies clés :

  • Azure / Microsoft Fabric : ADF, Event Hubs, ADLS Gen2, Synapse, Fabric, Purview, Azure Monitor, Log Analytics
  • AWS : Glue, Kinesis, Lambda, S3, Redshift
  • GCP : Dataflow, Pub/Sub, BigQuery
  • Databricks : Spark, Delta Lake, Autoloader
  • Snowflake, dbt, Airflow, Great Expectations, Soda, SQL, Python, OpenTelemetry

Processus de Recrutement :

  • 1er échange RH avec Jérôme LEONELLI ou Elise FORTEGUERRE : échange sur vos souhaits et motivation, présentation de l’entreprise (durée : 30mn, Teams ou présentiel)
  • 2ème échange avec un opérationnel (durée : 1h, Teams ou présentiel)
  • 3ème échange avec un opérationnel (durée : 1h, Teams ou présentiel)
  • 4ème échange avec un Partner : projection sur votre rôle au sein de Onepoint (durée : 1h, Teams ou présentiel)

Onepoint en quelques mots
Créée en 2002 par David Layani, Onepoint est devenu en un peu plus de 20 ans l’un des acteurs majeurs du conseil et de la tech, notre collectif compte aujourd’hui plus de 4000 talents en France (Aix-en-Provence, Bordeaux, Lyon, Nantes, Paris, Rennes, Strasbourg et Toulouse) et dans le monde (Australie, Belgique, Canada, Etats-Unis, Angleterre, Malaisie, Maroc et Singapour). Notre chiffre d’affaires a été multiplié par dix en 10 ans, atteignant plus de 500 millions d’euros, et ambitionne le milliard d’euros d’ici 4 ans.

Notre promesse
Un parcours sur mesure pour tous nos talents
Pour accompagner l’évolution de tous nos talents, nous proposons deux axes d’évolution complémentaires leur permettant de s’épanouir et de bénéficier des mêmes avantages : le business management et l’expertise.

L’école Onepoint pour développer vos expertises
Avec plus de 200 formations à son catalogue, notre école permet à nos talents de se former tout au long de leur vie professionnelle, en articulant travail et apprentissage.

Nos lieux de vie
Nos espaces de travail sont situés au cœur des territoires, au plus près de de nos talents de concilier projet de vie et projet professionnel.

Et bien plus encore ...
Notre écosystème permet de pratiquer de nombreuses activités loisirs et sportives (musique, photo, e-sport, yoga…) avec un cadre de travail flexible grâce à une charte de télétravail, et de prendre soin de tous (cabinet médical en libre accès, groupes de travail dédiés aux enjeux « égalité et pluralité » …).

Comme des images valent mieux que mille mots, découvrez notre ADN à travers cette vidéo : Nous sommes Onepoint✨

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

Ingénieur en Automatisation et Systèmes Robotisés (H/F) - Canada

EIC - Experience Internship Canada

Parisfulltime

Apply on the employer's site

Role description

EIC est le leader du placement de stagiaires internationaux au Canada. Notre mission est d’accompagner les étudiants internationaux afin de leur offrir une expérience de stage enrichissante en Amérique du Nord dans des conditions optimales !

Pourquoi choisir Experience Internship Canada ?
🌎

✅ Plus de 200 entreprises partenaires

✅ 17 secteurs d’activité : Finance, Marketing, Informatique & Technologie, Immobilier, Ressources humaines, ONG, Ingénierie, Data Science, Sports…

✅ Leader sur le marché des stages internationaux au Canada (Montréal, Toronto, Vancouver,...)

Ce que notre programme de stage offre aux étudiants :

  • Un stage professionnel dans ton domaine
  • Hébergement en centre-ville (salle de sport, piscine,...)
  • Assistance dans les démarches de visa
  • Soutien local 24h/24 et 7j/7
  • Coaching individuel et test de personnalité MBTI
  • Accueil à l'aéroport
  • Formation professionnelle en entreprise
  • Accès à un réseau international de contacts
  • Événements culturels et sociaux (afterwork, match de hockey,...)
  • Cours de langues (anglais ou français - optionnel)

Les avantages d'un stage à l'international :

  • Valoriser son CV avec une expérience internationale reconnue par les employeurs
  • Développer de nouvelles compétences professionnelles
  • Améliorer son anglais et travailler dans un environnement multiculturel
  • Se démarquer sur le marché de l’emploi grâce à une immersion professionnelle à l’étranger
  • Avoir des responsabilités concrètes
  • Bénéficier d'un accompagnement personnalisé favorisant leur développement personnel et professionnel
  • Elargir son réseau professionnel à l’échelle internationale
  • Découvrir la culture nord-américaine

Consulter notre site web pour plus d'informations : https://www.eic\-canada.ca/

Contexte (Stage au Canada)

Ce stage vous permettra de participer à des projets de conception et d’optimisation de systèmes robotisés, tout en vous familiarisant avec les technologies modernes de robotique. Vous aurez l’opportunité d’évoluer dans un environnement dynamique où vous travaillerez sur des applications concrètes et innovantes en robotique et automatisation.

🛠 Missions principales

  • Conception et programmation des systèmes robotisés
    : Participer à la conception et à la programmation des systèmes robotisés en utilisant des outils comme
    ROS
    (Robot Operating System),
    Python
    , ou
    C++
    pour l’automatisation de processus.
  • Optimisation des processus automatisés
    : Aider à l'optimisation des systèmes automatisés pour améliorer l'efficacité, la rapidité et la précision des robots dans des environnements industriels.
  • Tests et validation des robots
    : Contribuer à la réalisation de tests sur les robots, en analysant leurs performances et en ajustant les paramètres pour garantir leur bon fonctionnement dans des conditions réelles.
  • Gestion des capteurs et actionneurs
    : Participer à l'intégration et à la gestion des capteurs (caméras, LiDAR, capteurs de proximité) et des actionneurs (moteurs, servomoteurs) utilisés dans les systèmes robotisés.
  • Veille technologique et recherche d’innovations
    : Se tenir informé(e) des dernières tendances et technologies en automatisation et robotique pour proposer des solutions innovantes aux projets en cours.
  • Support à l’intégration des robots dans les chaînes de production
    : Aider à l’intégration des robots dans des processus de production automatisés, en collaborant avec les équipes techniques pour assurer une mise en œuvre efficace.

Compétences recherchées :

  • Automatisation et Robotique
    : Expérience avec des systèmes automatisés et logiciels de simulation comme
    MATLAB
    ,
    SolidWorks
    ,
    ANSYS
    .
  • Programmation et Contrôle
    : Langages
    C
    ,
    C++
    , et
    MATLAB
    pour le développement de systèmes robotiques et d’automatisation.
  • Multidisciplinarité
    : Capacité à travailler avec des ingénieurs de diverses spécialités pour créer des solutions globales et performantes.

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

Site Reliability Engineer, Mistral Cloud

Mistral

Parisfulltime

Apply on the employer's site

Role description

About Mistral
Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector—co-creating customized AI systems that they can run on their terms.

We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited.

The Role
As a
Site Reliability Engineer (SRE)
on the
Cloud Platform
team, you will shape the reliability, scalability, and performance of our Cloud platform and customer-facing applications. You’ll work closely with software engineers and product teams to ensure our systems meet and exceed the expectations of both internal and external customers.

This role is critical in maintaining the stability and efficiency of our infrastructure, enabling seamless experiences for users and developers. Your expertise will directly impact the robustness of our AI platform, ensuring it operates at scale with minimal downtime.

What You Will Do

  • Design, build, and maintain scalable, highly available, and fault-tolerant infrastructures to support our Cloud platform.
  • Operate systems and troubleshoot issues in production environments, including interrupts, on-call responses, and infrastructure scaling.
  • Implement and improve monitoring, alerting, and incident response systems to minimize downtime and optimize performance.
  • Develop and maintain workflows and tools for CI/CD, containerization, orchestration, monitoring, and logging.
  • Participate in on-call rotations to respond to incidents and perform root cause analysis to prevent recurrence.
  • Drive continuous improvement in infrastructure automation, deployment, and orchestration.
  • Collaborate with software engineers to enable safe and reproducible model-training experiments.
  • Build and enhance a cloud platform that abstracts infrastructure complexities for science and engineering teams.
  • Design and develop new workflows, tooling, and automation to improve system reliability, availability, and performance.
  • Ensure infrastructure adheres to security best practices and compliance requirements in collaboration with the security team.
  • Document processes and procedures to ensure consistency and knowledge sharing across the team.

What We're Looking For

  • A Master’s degree in Computer Science, Engineering, or a related field.
  • 5+ years of experience in a DevOps or SRE role, with strong expertise in bare metal infrastructure and distributed systems.
  • Hands-on experience with site reliability issues, including root cause analysis, in-production troubleshooting, and on-call rotations.
  • Proficiency in working with reliability KPIs, such as observability, alerting, and SLAs.
  • Experience with CI/CD, containerization, and orchestration tools like Docker and Kubernetes.
  • Knowledge of monitoring, logging, alerting, and observability tools such as Prometheus, Grafana, ELK Stack, or Datadog.
  • Familiarity with infrastructure-as-code tools like Terraform or CloudFormation.
  • Proficiency in scripting languages (Python, Go, Bash) and a strong understanding of software development best practices.
  • Solid grasp of networking, security, and system administration concepts.
  • Excellent problem-solving and communication skills, with the ability to work effectively in a collaborative environment.
  • Experience in an AI/ML environment, high-performance computing (HPC) systems, or modern AI-oriented solutions (e.g., Fluidstack, Coreweave, Vast) is a plus.

What we offer
We offer a comprehensive benefits package designed to support your well-being, growth, and work-life balance. Benefits vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks.

For the most up-to-date details on benefits available in your location, please refer to our Benefits page.

Privacy Policy
Your privacy matters to us. You can learn more about how we handle your personal data in our Applicant Privacy Policy.

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

Site Reliability Engineer

Mistral

Parisfulltime

Apply on the employer's site

Role description

About Mistral
Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector—co-creating customized AI systems that they can run on their terms.

We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited.

The Role
As a
Site Reliability Engineer (SRE)
on the
Platform
team, you will shape the reliability, scalability, and performance of our platform and customer-facing applications. You’ll work closely with software engineers and research teams to ensure our systems meet and exceed the expectations of both internal and external customers.

This role balances day-to-day operations on production systems with long-term software engineering improvements. Your work will reduce operational toil, foster reliability, and ensure high availability for our web services, inference environments, and ML workloads. You’ll enable seamless replication of work environments across multiple HPC clusters, directly impacting the stability and efficiency of our AI platform.

What You Will Do

  • Design, build, and maintain scalable, highly available, and fault-tolerant infrastructures to support web services and ML workloads.
  • Ensure our platform, inference, and model training environments are always highly available and enable seamless replication across HPC clusters.
  • Operate systems and troubleshoot issues in production, including interrupts, on-call responses, and infrastructure scaling.
  • Implement and improve monitoring, alerting, and incident response systems to minimize downtime and optimize performance.
  • Develop and maintain workflows and tools for CI/CD, containerization, orchestration, monitoring, and logging.
  • Participate in on-call rotations to respond to incidents and perform root cause analysis.
  • Drive continuous improvement in infrastructure automation, deployment, and orchestration using tools like Kubernetes, Flux, and Terraform.
  • Collaborate with AI/ML researchers to enable safe and reproducible model-training experiments.
  • Build a cloud-agnostic platform that abstracts infrastructure complexities for science and engineering teams.
  • Design and develop new workflows, tooling, and automation to improve system reliability, availability, and performance.
  • Work with the security team to ensure infrastructure adheres to best practices and compliance requirements.
  • Document processes and procedures to ensure consistency and knowledge sharing across the team.

What We're Looking For

  • A Master’s degree in Computer Science, Engineering, or a related field.
  • 7+ years of experience in a DevOps or SRE role, with strong expertise in cloud computing and distributed systems.
  • Hands-on experience with site reliability issues, including root cause analysis, in-production troubleshooting, and on-call rotations.
  • Proficiency in working with reliability KPIs, such as observability, alerting, and SLAs.
  • Experience with CI/CD, containerization, and orchestration tools like Docker and Kubernetes.
  • Knowledge of monitoring, logging, alerting, and observability tools such as Prometheus, Grafana, ELK Stack, or Datadog.
  • Familiarity with infrastructure-as-code tools like Terraform or CloudFormation.
  • Proficiency in scripting languages (Python, Go, Bash) and a strong understanding of software development best practices.
  • Solid grasp of networking, security, and system administration concepts.
  • Excellent problem-solving and communication skills, with the ability to work effectively in a collaborative environment.
  • Experience in an AI/ML environment, high-performance computing (HPC) systems, or modern AI-oriented solutions (e.g., Fluidstack, Coreweave, Vast) is a plus.

What we offer
We offer a comprehensive benefits package designed to support your well-being, growth, and work-life balance. Benefits vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks.

For the most up-to-date details on benefits available in your location, please refer to our Benefits page.

Privacy Policy
Your privacy matters to us. You can learn more about how we handle your personal data in our Applicant Privacy Policy.

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

Senior Site Reliability Engineer (Platform Engineering)

Criteo

Parisfulltime

Apply on the employer's site

Role description

What You'll Do:

Within the Platform department of R&D, the Product Reliability Engineering (PRE) group acts as the bridge between Product Engineering, Platform Engineering and Infrastructure. The Platform PRE group comprises eight teams helping R&D design, build, and operate large-scale distributed systems reliably and efficiently. The common objective of the PRE teams is to build the most reliable platform in AdTech.

As a Senior Site Reliability Engineer within the PRE-Platform team, you will help Platform teams improve the reliability, scalability, and operability of their services and CI/CD pipelines. You will develop and maintain tooling and libraries (Python, Jenkins, Chef), support observability and performance initiatives, and lead technical migrations across infrastructure and core dependencies. You will also participate in an on-call rotation to help ensure production stability.

What You’ll Do

  • Own and improve the full lifecycle of production services — from design and deployment to operation and continuous improvement.
  • Partner with development teams before launch through system design reviews, platform and framework development, capacity planning, and production readiness assessments.
  • Improve reliability, scalability, and performance by automating operations and driving infrastructure and platform enhancements.
  • Maintain and optimize live systems through monitoring, observability, performance analysis, and incident management.
  • Participate in incident response and contribute to a culture of blameless postmortems and continuous learning.
  • Stack: C#, Java, Scala, Python, Go, Prometheus, Grafana, Kibana, Linux, Kubernetes, Mesos, and more.

Who You Are:

  • Strong software engineering experience with at least one modern programming language.
  • 5+ years of experience in Site Reliability Engineering, Platform Engineering, Software Engineering, or DevOps roles.
  • Experience designing, operating, and troubleshooting large-scale distributed systems in production environments.
  • Solid understanding of systems engineering fundamentals, including compute, networking, storage, and observability.
  • Hands-on experience debugging production issues, optimizing system performance, and automating operational workflows.
  • Ability to write clean, maintainable, and reliable code used in production systems and internal platforms.
  • Strong analytical and problem-solving skills, with the ability to collaborate effectively across engineering teams.
  • Curiosity, ownership, and a pragmatic mindset toward improving reliability and developer productivity.

Take a look at our R&D blog on Medium for access and insight into our engineering culture and achievements.

We acknowledge that many candidates may not meet every single role requirement listed above. If your experience looks a little different from our requirements but you believe that you can still bring value to the role, we’d love to see your application!

Who We Are:

We’re Criteo, the Commerce Intelligence Platform. Criteo helps businesses turn shopper signals into commerce outcomes while delivering more relevant experiences for shoppers. We use proprietary commerce intelligence and AI decisioning to drive relevance for shoppers and performance for businesses.

At Criteo, our culture is as unique as it is diverse. From our offices across the globe or from the comfort of home, our 3,600 Criteos collaborate together to build an open, impactful, and forward-thinking environment.

We foster a workplace where everyone is valued, and employment decisions are based solely on skills, qualifications, and business needs—never on non-job-related factors or legally protected characteristics.

What We Offer:

🏢 Ways of working – Our hybrid model blends home with in-office experiences, making space for both.

📈 Grow with us – Learning, mentorship & career development programs.

💪 Your wellbeing matters – Health benefits, wellness perks & mental health support.

🤝 A team that cares – Diverse, inclusive, and globally connected.

💸 Fair pay & perks – Attractive salary, with performance-based rewards and family-friendly policies, plus the potential for equity depending on role and level.

Additional benefits may vary depending on the country where you work and the nature of your employment with Criteo.

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

79 more openings in this category and country

Consultant Data Ops – F/HOnepoint · France

Apply on the employer's site