Skip to content
Site Reliability Engineer

The week's list

Every role like this one, in one letter

You are reading one posting. There are hundreds like it on the board, and new ones every week. Pick what you want, leave an email, and the list comes to you — no searching, no coming back here.

Counting what came out this past week…

The first letter arrives right away, then one a week. Unsubscribe in one click from any letter — the address goes nowhere else.

Site Reliability Engineer

Ingénieur production OPS H/F

Crédit Agricole Group Infrastructure Platform

Guyancourt

Apply on the employer's site

Role description

Vous êtes
Ingénieur production OPS ?
Devenez partenaire des grandes évolutions chez CA-GIP !

Qui sommes-nous ?
Gérant 80% de la production informatique du groupe Crédit Agricole, nous formons une véritable communauté d’experts engagés. Nous mobilisons nos expertises afin de développer des plateformes et services adaptés aux nouvelles pratiques du digital tout en maintenant un haut niveau de confidentialité, et de sécurité.

C
ONTEXTE :
Au sein du cluster Services Financiers Spécialisés, la Tribu CAL&F Factoring a en charge la gestion de l’exploitation des infrastructures du client CAL&F. L’équipe assure le suivi du traitement des incidents, requêtes, changements et problèmes du périmètre.

Rattaché au leader de la Tribu, le collaborateur travaillera en relation étroite avec ses collègues, afin d’assurer l'exploitation et le maintien en conditions opérationnelles (MCO) des infrastructures.

Dans le cadre de ses missions il sera en relation étroite avec l’ensemble des équipes du cluster, ainsi qu'avec les équipes techniques CA-GIP (socles techniques) et les équipes de la DSI du client.

MISSIONS :
L’ingénieur Production préconise et met en œuvre les solutions méthodologiques et techniques permettant d’optimiser la production informatique.

Il assure la coordination et le contrôle de la qualité de l’intégration d’une solution.

Il permet de maintenir le legacy (MCO) mais ses compétences au sein du monde DevOps sont croissantes.

Localisé dans les clusters, il doit allier bonne connaissance du patrimoine applicatif géré et un vernis technologique pour assurer le lien avec les experts dans les socles.

Il a notamment en charge de :

  • Analyser les besoins OS ;
  • Assurer les mises en production et l’intégration de solutions ;
  • Gérer le Run quotidien via une activité de contrôle, surveillance et de gestion des incidents ;
  • Contribuer aux projets (Build) sous le pilotage d’un Chef de Projet ;
  • Conseiller les Entités (MOA) et collaborer avec les socles ;
  • Rechercher l’optimisation de l’intégration applicative ;
  • Assurer la montée de nouvelles versions.

Compétences métiers :

  • Python (langage) ;
  • OS Windows ;
  • Linux

Softskills :

  • Autonomie ;
  • Authenticité ;
  • Assertivité ;
  • Anticipation ;
  • Capacité d'analyse et de synthèse

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

Ingénieur système RUN (H/F)

Tenexa

Colombessenior

Apply on the employer's site

Role description

Chez Tenexa, ESN spécialisée dans l’IT, nous accompagnons nos clients sur des projets à forte valeur ajoutée en plaçant l’expertise et l’humain au cœur de nos interventions.

Dans le cadre de notre développement, nous recherchons un(e)
Ingénieur Système RUN
(H/F) pour rejoindre notre équipe à
Colombes
.

Vos missions.

Rattaché(e) à la Direction Technique et Innovation, vous intégrez l’équipe OPS en charge du RUN du Cloud Horizon.

En lien avec les équipes internes, les centres de services d’infogérance et certains clients grands comptes, vous contribuez à garantir la disponibilité, la performance et la sécurité des services Cloud.

À ce titre, vous serez amené(e) à :

  • Assurer le maintien en conditions opérationnelles du système d’information supportant les services de l’offre Cloud Horizon.
  • Traiter les demandes et incidents complexes escaladés par les équipes Front Office.
  • Instruire et mettre en œuvre les changements d’infrastructure d’origine client ou interne.
  • Piloter les analyses de causes racines et contribuer à la résolution des problèmes dans une démarche d’amélioration continue.
  • Participer aux transitions de projets dans le cadre des activités « Build to Run ».
  • Réaliser des transferts de compétences et accompagner les équipes d’exploitation sur les procédures techniques.
  • Contribuer à l’automatisation des tâches récurrentes via le scripting et les outils d’automatisation.
  • Participer aux astreintes et interventions en heures non ouvrées lorsque la criticité des services l’exige.

Vous disposez d’environ 5 ans d’expérience sur un poste d’administrateur ou d’ingénieur systèmes, idéalement au sein d’environnements cloud ou d’infrastructures complexes.

Compétences techniques requises

Systèmes

  • Windows Server (Active Directory, PKI, RDS, WSFC…)
  • Linux (Red Hat ou équivalent)
  • Administration et exploitation de serveurs physiques

Virtualisation

  • VMware ESXi / vCenter

Sauvegarde & Stockage

  • Veeam Backup & Replication
  • SAN (iSCSI, Fibre Channel)

Conteneurisation

  • Docker / Podman
  • Kubernetes apprécié

Automatisation

  • PowerShell
  • Bash ou Python
  • Ansible

Bases de données

  • SQL Server
  • MySQL / MariaDB

Gestion de parc et patch management

  • Ivanti
  • Red Hat Satellite / Uyuni

Sécurité

  • Bonnes pratiques de sécurité et référentiels ISO 27001
  • Gestion des certificats et infrastructures PKI
  • Sensibilité aux enjeux de conformité et de cybersécurité dans un contexte cloud

Compétences appréciées

  • Interventions en datacenter
  • HPE Morpheus
  • NetVault ou Commvault
  • eObserve / ServiceNav
  • Connaissances réseaux et firewall

Qualités personnelles

  • Curiosité technique et envie d’apprendre
  • Force de proposition et sens de l’amélioration continue
  • Esprit d’équipe et capacité à partager ses connaissances
  • Rigueur, organisation et esprit d’analyse
  • Bonne gestion des priorités et des situations critiques

Ce que nous proposons

  • Des missions variées et adaptées à votre profil
  • Un environnement technique stimulant autour des technologies Cloud et Infrastructure
  • Un accompagnement dans votre montée en compétences
  • Une ambiance d’équipe conviviale et un suivi régulier
  • L’opportunité de contribuer à l’évolution et à la fiabilité de services critiques

Informations complémentaires

  • Poste basé à Colombes (92)
  • Démarrage à compter de juillet 2026
  • Environ 10 % de l’activité réalisée en heures non ouvrées ou le week-end
  • Astreintes régulières à prévoir

Envie d’en savoir plus ? Postulez, et discutons ensemble de votre projet professionnel.

L'entreprise en quelques mots
Rejoindre TENEXA…
Depuis 1985, Tenexa s’est imposé comme un acteur incontournable des services numériques. Le groupe accompagne les entreprises dans leur transformation digitale en offrant des solutions sur-mesure et innovante en Cloud, Cybersécurité et Infogérance IT, répondant aux enjeux des PME et ETI d’aujourd’hui et de demain.

Une expertise à 360°
Tenexa couvre l’ensemble du cycle de vie des services IT, du DataCenter jusqu’à l’utilisateur final, garantissant performance et sécurité aux organisations de toutes tailles.

Sa stratégie repose sur une innovation continue et une excellence opérationnelle éprouvée, au service de plus de 600 clients, dont des ETI et de grands groupes issus de divers secteurs.

Une ambition soutenue par un actionnaire de référence
Le Groupe Tenexa bénéficie du soutien de HLD, société d’investissement à capitaux permanents, dédiée aux entreprises européennes à fort potentiel. Cet appui stratégique renforce la vision à long terme de Tenexa et accélère son développement sur des marchés clés.

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

Site Reliability Engineer (m/f/d)

Allianz Partners

Saint-Ouenfulltime

Apply on the employer's site

Role description

Key Responsibilities
As a Site Reliability Engineer within Advanced Analytics (DA3) in the Chief Data & AI Office at Allianz Partners, you will join the platform engineering team to own the reliability and operational health of the central engineering platform.

You will define and maintain service level objectives, drive incident response at the infrastructure layer, and systematically eliminate operational toil through automation.

You will work closely with Platform Engineers, Security Engineers, and incident-response leads to ensure the platform meets its reliability commitments across production workloads spanning AI services, Java APIs, and frontend applications.

Through this role, you will have the main following responsibilities:

  • Define, instrument, and maintain SLOs and SLIs for platform components; own error budget tracking and produce regular reliability reports for senior leadership.
  • Serve on the on-call rotation as the infrastructure escalation tier; lead incident response for cluster-level, network-level, and storage failures; chair blameless post-incident reviews.
  • Implement and operate Kubernetes infrastructure (AKS): cluster lifecycle management, networking, resource quotas, autoscaling configuration, and multi-tenancy patterns across product team namespaces.
  • Develop Infrastructure as Code (Terraform) to provision and manage Azure resources with consistency, auditability, and repeatable rollback capability.
  • Build and maintain observability infrastructure: Prometheus, Grafana, Azure Monitor, and Application Insights; own alerting rules, dashboards, and distributed tracing coverage across platform components.
  • Perform capacity planning and cost-aware resource management: right-size node pools, tune vertical and horizontal pod autoscalers, and identify resource waste across namespaces.
  • Identify and eliminate toil: automate repetitive operational tasks through scripting and tooling; measure and track toil reduction over time.
  • Maintain platform reliability procedures: rolling upgrades, backup and recovery testing, disaster recovery runbooks, and change freeze coordination.
  • Contribute to CI/CD pipelines and GitOps tooling (GitHub Actions, ArgoCD) from a reliability and deployment safety perspective; work with platform engineering on release gates and rollback mechanisms.
  • Collaborate with incident-response leads on incident SLA targets and operational procedures; work with Security Engineers on infrastructure hardening and vulnerability remediation.

What You Bring

  • 5+ years professional experience in site reliability engineering, DevOps, or platform engineering roles.
  • Strong Kubernetes experience: cluster operations, networking (Ingress, network policies), storage, autoscaling, and hands-on troubleshooting across production environments.
  • Solid Infrastructure as Code experience with Terraform; familiarity with Bicep or ARM templates is a plus.
  • Production experience with Azure cloud services: AKS, ACR, Key Vault, Azure Monitor, Application Insights, Virtual Networks, and Private Endpoints.
  • Strong observability experience: Prometheus, Grafana, centralized logging, alerting configuration, and distributed tracing instrumentation.
  • Working knowledge of SLO/SLI methodology: error budget principles, reliability target setting, and capacity planning.
  • Structured incident management experience: on-call ownership, blameless post-incident review, and runbook authorship.
  • Scripting and automation proficiency in Python or bash for toil elimination and operational tooling.
  • Strong CI/CD experience: GitHub Actions and ArgoCD or equivalent GitOps tooling.

Ways of Working

  • Comfortable in agile, iterative delivery environments with personal ownership and accountability for platform reliability.
  • Clear communicator across global, cross-functional stakeholders; able to translate technical reliability metrics into business impact for non-technical audiences.
  • Proactive learner with pragmatic adoption of AI-assisted developer tools (e.g., GitHub Copilot, Claude Code) to improve automation coverage and delivery velocity.

Nice to Have

  • Kubernetes certifications: CKA or CKAD.
  • Experience supporting AI or ML infrastructure workloads: GPU scheduling, model serving platforms, or inference pipeline operations.
  • Exposure to chaos engineering practices and fault injection testing.
  • FinOps experience: reserved capacity planning, resource right-sizing programs, and cost attribution per team or workload.
  • Service mesh experience (Istio, Linkerd) for traffic management and reliability patterns.
  • Experience in regulated industries (insurance, finance, healthcare) where auditability, change traceability, and secure-by-default operations are standard practice.

How We Hire
Allianz Partners does not accept unsolicited CV’s or approaches from agencies. We only work with partners on our approved supplier lists, under contract. Any unsolicited submission will not be considered.

What We Offer
Our employees play an integral part in our success as a business. We appreciate that each of our employees are unique and have unique needs, ambitions and we enjoy being a part of their journey. We are there to empower and encourage you with your personal and professional development ensuring that you take control by offering a large variety of courses and targeted development programs.

All that in a global environment where international mobility and career progression are encouraged. Caring for your health and wellbeing is key priority for us. This is why we build Work Well programs to providing you with peace of mind and give the flexibility in planning and arranging for a better work-life balance.

90377 | Data & AI | Professional | Allianz Partners | Full-Time | Permanent

Allianz Group is one of the most trusted insurance and asset management companies in the world. Caring for our employees, their ambitions, dreams and challenges, is what makes us a unique employer. Together we can build an environment where everyone feels empowered and has the confidence to explore, to grow and to shape a better future for our customers and the world around us.

At Allianz, we stand for unity: we believe that a united world is a more prosperous world, and we are dedicated to consistently advocating for equal opportunities for all. And the foundation for this is our inclusive workplace, where people and performance both matter, and nurtures a culture grounded in integrity, fairness, inclusion and trust.

We therefore welcome applications regardless of ethnicity or cultural background, age, gender, nationality, religion, social class, disability or sexual orientation, or any other characteristics protected under applicable local laws and regulations.

Join us. Let's care for tomorrow.

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

Ingénieur en Automatisation et Systèmes Robotisés (H/F) - Canada

EIC - Experience Internship Canada

Parisfulltime

Apply on the employer's site

Role description

EIC est le leader du placement de stagiaires internationaux au Canada. Notre mission est d’accompagner les étudiants internationaux afin de leur offrir une expérience de stage enrichissante en Amérique du Nord dans des conditions optimales !

Pourquoi choisir Experience Internship Canada ?
🌎

✅ Plus de 200 entreprises partenaires

✅ 17 secteurs d’activité : Finance, Marketing, Informatique & Technologie, Immobilier, Ressources humaines, ONG, Ingénierie, Data Science, Sports…

✅ Leader sur le marché des stages internationaux au Canada (Montréal, Toronto, Vancouver,...)

Ce que notre programme de stage offre aux étudiants :

  • Un stage professionnel dans ton domaine
  • Hébergement en centre-ville (salle de sport, piscine,...)
  • Assistance dans les démarches de visa
  • Soutien local 24h/24 et 7j/7
  • Coaching individuel et test de personnalité MBTI
  • Accueil à l'aéroport
  • Formation professionnelle en entreprise
  • Accès à un réseau international de contacts
  • Événements culturels et sociaux (afterwork, match de hockey,...)
  • Cours de langues (anglais ou français - optionnel)

Les avantages d'un stage à l'international :

  • Valoriser son CV avec une expérience internationale reconnue par les employeurs
  • Développer de nouvelles compétences professionnelles
  • Améliorer son anglais et travailler dans un environnement multiculturel
  • Se démarquer sur le marché de l’emploi grâce à une immersion professionnelle à l’étranger
  • Avoir des responsabilités concrètes
  • Bénéficier d'un accompagnement personnalisé favorisant leur développement personnel et professionnel
  • Elargir son réseau professionnel à l’échelle internationale
  • Découvrir la culture nord-américaine

Consulter notre site web pour plus d'informations : https://www.eic\-canada.ca/

Contexte (Stage au Canada)

Ce stage vous permettra de participer à des projets de conception et d’optimisation de systèmes robotisés, tout en vous familiarisant avec les technologies modernes de robotique. Vous aurez l’opportunité d’évoluer dans un environnement dynamique où vous travaillerez sur des applications concrètes et innovantes en robotique et automatisation.

🛠 Missions principales

  • Conception et programmation des systèmes robotisés
    : Participer à la conception et à la programmation des systèmes robotisés en utilisant des outils comme
    ROS
    (Robot Operating System),
    Python
    , ou
    C++
    pour l’automatisation de processus.
  • Optimisation des processus automatisés
    : Aider à l'optimisation des systèmes automatisés pour améliorer l'efficacité, la rapidité et la précision des robots dans des environnements industriels.
  • Tests et validation des robots
    : Contribuer à la réalisation de tests sur les robots, en analysant leurs performances et en ajustant les paramètres pour garantir leur bon fonctionnement dans des conditions réelles.
  • Gestion des capteurs et actionneurs
    : Participer à l'intégration et à la gestion des capteurs (caméras, LiDAR, capteurs de proximité) et des actionneurs (moteurs, servomoteurs) utilisés dans les systèmes robotisés.
  • Veille technologique et recherche d’innovations
    : Se tenir informé(e) des dernières tendances et technologies en automatisation et robotique pour proposer des solutions innovantes aux projets en cours.
  • Support à l’intégration des robots dans les chaînes de production
    : Aider à l’intégration des robots dans des processus de production automatisés, en collaborant avec les équipes techniques pour assurer une mise en œuvre efficace.

Compétences recherchées :

  • Automatisation et Robotique
    : Expérience avec des systèmes automatisés et logiciels de simulation comme
    MATLAB
    ,
    SolidWorks
    ,
    ANSYS
    .
  • Programmation et Contrôle
    : Langages
    C
    ,
    C++
    , et
    MATLAB
    pour le développement de systèmes robotiques et d’automatisation.
  • Multidisciplinarité
    : Capacité à travailler avec des ingénieurs de diverses spécialités pour créer des solutions globales et performantes.

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

Site Reliability Engineer, Mistral Cloud

Mistral

Parisfulltime

Apply on the employer's site

Role description

About Mistral
Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector—co-creating customized AI systems that they can run on their terms.

We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited.

The Role
As a
Site Reliability Engineer (SRE)
on the
Cloud Platform
team, you will shape the reliability, scalability, and performance of our Cloud platform and customer-facing applications. You’ll work closely with software engineers and product teams to ensure our systems meet and exceed the expectations of both internal and external customers.

This role is critical in maintaining the stability and efficiency of our infrastructure, enabling seamless experiences for users and developers. Your expertise will directly impact the robustness of our AI platform, ensuring it operates at scale with minimal downtime.

What You Will Do

  • Design, build, and maintain scalable, highly available, and fault-tolerant infrastructures to support our Cloud platform.
  • Operate systems and troubleshoot issues in production environments, including interrupts, on-call responses, and infrastructure scaling.
  • Implement and improve monitoring, alerting, and incident response systems to minimize downtime and optimize performance.
  • Develop and maintain workflows and tools for CI/CD, containerization, orchestration, monitoring, and logging.
  • Participate in on-call rotations to respond to incidents and perform root cause analysis to prevent recurrence.
  • Drive continuous improvement in infrastructure automation, deployment, and orchestration.
  • Collaborate with software engineers to enable safe and reproducible model-training experiments.
  • Build and enhance a cloud platform that abstracts infrastructure complexities for science and engineering teams.
  • Design and develop new workflows, tooling, and automation to improve system reliability, availability, and performance.
  • Ensure infrastructure adheres to security best practices and compliance requirements in collaboration with the security team.
  • Document processes and procedures to ensure consistency and knowledge sharing across the team.

What We're Looking For

  • A Master’s degree in Computer Science, Engineering, or a related field.
  • 5+ years of experience in a DevOps or SRE role, with strong expertise in bare metal infrastructure and distributed systems.
  • Hands-on experience with site reliability issues, including root cause analysis, in-production troubleshooting, and on-call rotations.
  • Proficiency in working with reliability KPIs, such as observability, alerting, and SLAs.
  • Experience with CI/CD, containerization, and orchestration tools like Docker and Kubernetes.
  • Knowledge of monitoring, logging, alerting, and observability tools such as Prometheus, Grafana, ELK Stack, or Datadog.
  • Familiarity with infrastructure-as-code tools like Terraform or CloudFormation.
  • Proficiency in scripting languages (Python, Go, Bash) and a strong understanding of software development best practices.
  • Solid grasp of networking, security, and system administration concepts.
  • Excellent problem-solving and communication skills, with the ability to work effectively in a collaborative environment.
  • Experience in an AI/ML environment, high-performance computing (HPC) systems, or modern AI-oriented solutions (e.g., Fluidstack, Coreweave, Vast) is a plus.

What we offer
We offer a comprehensive benefits package designed to support your well-being, growth, and work-life balance. Benefits vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks.

For the most up-to-date details on benefits available in your location, please refer to our Benefits page.

Privacy Policy
Your privacy matters to us. You can learn more about how we handle your personal data in our Applicant Privacy Policy.

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

52 more openings in this category and country

Ingénieur production OPS H/FCrédit Agricole Group Infrastructure Platform · France

Apply on the employer's site