Skip to content
Site Reliability Engineer

Список недели

Все такие вакансии — одним письмом

Сейчас вы читаете одну вакансию. Таких же на витрине сотни, и каждую неделю выходят новые. Выберите, что присылать, оставьте почту — список придёт сам, без поисков и без возвращения сюда.

Считаем, сколько вышло за прошлую неделю…

Первое письмо приходит сразу, дальше — раз в неделю. Отписка в один клик из любого письма, адрес больше никуда не уходит.

Site Reliability Engineer

Site Reliability Engineer - Storage

Proton

Paris

Откликнуться на сайте работодателя

Описание вакансии

Join Proton and build a better internet where privacy is the default

Proton was founded in 2014 by scientists from CERN on a simple truth:
privacy is a fundamental human right
. Since then, we've built the world's largest encrypted email service (Proton Mail) and expanded into Proton VPN, Proton Drive, Proton Pass, and Proton Calendar — tools used by millions globally to protect their freedom, fight censorship, and keep their data safe. In some situations, Proton has literally helped save lives.

We are profitable, independent (no VC control), and selectively hire from the top :1% of applicants. Our 700+ team members across 50+ countries come from leading organizations and elite academic backgrounds. We move fast, keep hierarchy light, and prioritize impact over optics. If you want to do meaningful work with exceptionally high-caliber people, this is it. Check our open-source projects here.

Purpose of the role:

Storage is at the very center of everything we do. Whatever needs to be persisted eventually uses our Storage infrastructure. For this we run our infrastructure predominantly on-premise in multiple data centers. In this role you will join the Storage SRE team that is responsible to automate, scale, enhance and run our Storage infrastructure. The office locations we hire into for this role are Geneva/Switzerland or Paris/France.

What you will do:

  • Design, build, and operate resilient Storage infrastructure for Proton’s most critical production services
  • Develop and maintain automation tools to enhance deployment processes, configuration management, and service reliability
  • Improve monitoring, alerting, and incident response to maintain high availability and rapid issue resolution
  • Optimize and scale Proton’s infrastructure
  • Collaborate with product and security teams to implement best practices for privacy and data protection at every layer of the stack
  • Document and share knowledge to foster a culture of learning and continuous improvement within the team and beyond
  • Work closely with a group of Back-End Engineers and Architects to build resilient, massively-scalable solutions for millions of people to rely on
  • Document back-end capabilities and their proper use
  • Join the oncall rotation to operate the Storage service

Job requirements:

  • Experience building complex production systems
  • Operational experience, demonstrating the ability to diagnose and troubleshoot complex problems for critical systems
  • Experience with software development
  • Solid understanding of security best practices
  • High-level understanding of cryptography concepts

Bonus points for:

  • Experience with storage systems (e.g. Ceph, Seaweed or Tape)

Even if you don’t meet all the requirements listed above, but feel you could still be a great fit, please apply.

What We Offer

  • Office First: Collaboration is easier and more effective in person, which is why we have offices in Geneva, Zurich, Prague, Barcelona, Paris, London, Vilnius, Skopje, and Taipei. You can also enjoy working from home up to 30% of the time, while enjoying great company during our three core days in the office.
  • Technology: We provide all the devices and software you need to excel in your role, ensuring you have the best tools at your disposal to achieve your goals.
  • Food: Lunch and snacks are provided by Proton every day at our offices.
  • Transport: We will always support our employees with transport costs through subsidizing public transport, bike allowances, or parking spaces based on your office location.
  • Salary Range : Paris - Base Annual Gross - 50,000€ to 77,000€ - Final compensation will be determined based on the candidate's qualifications, skills, and previous experience.
  • Stock Options: At Proton, we are all owners of the company and you'll be offered stock options when you join us.
  • Flexible Working: You can define your own working hours as long as it works with team meetings.
  • Learning and Development: We are committed to your professional growth. Proton offers various learning opportunities, including training programs, conferences and events, and continual learning.
  • Employee Benefits: Comprehensive health insurance plans, competitive retirement savings options, generous vacation and leave policies, and wellness programs.
  • Work that Matters: Proton is a community-first organization, started with the support of a crowdfunding campaign and built with community input. To this day, Proton’s only source of revenue is user subscriptions. Over 100 million people trust and support Proton, and we put our users and community first in everything we do.

Our Commitment to Diversity and Inclusion
At Proton, we believe diversity drives innovation and strengthens our mission to provide privacy as a default for all. We are committed to fostering an inclusive environment where all individuals, regardless of race, ethnicity, gender, age, sexual orientation, physical ability, or socio-economic background, feel valued and empowered. We strive to create equal opportunities, promote open dialogue, and support continuous learning to ensure every voice is heard and respected.

If you need any extra support or reasonable adjustments during the hiring process, please let your talent partner know.

Candidate Privacy Notice
When you apply for a position, refer a candidate, or are considered for a role at Proton Technologies AG (Proton, we, us, or our), your information is stored in Greenhouse, in accordance with their Service Privacy Policy. This information is used to evaluate your suitability for the posted position. We also retain this information for consideration for future roles that you may apply for or that we believe may align with your background and skills.

If we no longer have a legitimate business need to process your information, we will either delete or anonymize it. Should you have any inquiries about how we use or manage your information, or if you wish to access, correct, or delete your data, please contact our privacy team at careers@proton.ch.

Proton does not accept unsolicited resumes from any sources other than directly from candidates. We will not pay a fee for any placement resulting from an unsolicited offer, even if the candidate is subsequently hired by Proton.
To learn more about our privacy policy, please visit
our privacy policy page
.

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

Site Reliability Engineer

Staff Site Reliability Engineer (x/f/m)

Doctolib

Parislead

Откликнуться на сайте работодателя

Описание вакансии

Your Impact

We are looking for a
Staff Site Reliability Engineer
to join our SRE team dedicated to platform reliability within Platform Engineering.

Your mission will be to act as a technical leader driving Doctolib's reliability and scalability at a European scale, ensuring our platform remains reliable, debuggable, and resilient across infrastructure, observability, and cross-cutting reliability initiatives. You will play a pivotal role in a team driving reliability standards across 170+ applications, contributing directly to supporting 520,000 health professionals and 90 million patients in their daily healthcare journey.

This role sits at the intersection of infrastructure, developer experience, and product engineering. You'll act as a technical leader and strategic partner to SREs, software engineers, and product teams, guiding decisions, mentoring engineers, and driving cross-cutting initiatives that elevate our operational maturity.

What you'll do

Your responsibilities include but are not limited to:

  • Lead large-scale cross-cutting reliability initiatives across the platform, spanning infrastructure automation, observability, and incident management
  • Identify and drive improvements to incident detection, response, and postmortem analysis capabilities
  • Define and evolve SLOs, error budgets, and alerting standards across multiple product teams
  • Take part in the on-call rotation, and actively contribute to improving our on-call experience by refining alerting, reducing noise, and ensuring actionable telemetry
  • Serve as a mentor and technical coach to senior engineers, helping elevate the craft of reliability engineering across the company
  • Influence strategic decisions by providing technical guidance to leadership and representing reliability engineering in architectural reviews and platform discussions
  • Partner with software engineering teams to embed reliability practices early in the development lifecycle

Who you are

Before you read on: if you don't have the exact profile described below, but you feel this job description matches your skill set, we still encourage you to apply.

You'll be a great fit if you:

  • Have extensive experience (8+ years) in SRE, platform engineering, or infrastructure roles within a large-scale, multi-team production environment
  • Have proven experience with cloud platforms such as AWS, GCP, or Azure
  • Have strong experience with containerization and orchestration technologies, Kubernetes is a must, its deployment and scaling strategies ecosystem
  • Have implemented and operated SLIs, SLOs, and error budgets in production
  • Have experience managing on-call rotations and leading incident response in high-stakes environments
  • Have a strong systems engineering background with fluency in at least one backend programming language (e.g., Go, Python, Ruby)
  • Have a proven ability to lead through influence: setting technical direction, driving consensus, and mentoring engineers across teams
  • Are comfortable balancing long-term architecture work with fast, iterative improvements
  • Have clear, concise communication skills, both written and verbal, with the ability to drive alignment in ambiguous environments
  • Partner with feature teams to accelerate their production readiness, providing hands-on guidance on reliability best practices, launch reviews, and operational standards before go-live
  • Are fluent in English

It would be fantastic if you:

  • Have deep expertise in observability tooling and architecture (logging, tracing, metrics)
  • Have experience designing and operating high-scale telemetry pipelines and working with developers to improve instrumentation quality
  • Appreciate working in regulated environments, healthcare, fintech, or similar; at Doctolib, data privacy and compliance are part of every engineering decision
  • Care about reliability enablement, golden paths, runbooks, shared libraries

Life at Doctolib Tech

  • Our solutions are built on a single fully cloud-native platform that supports web and mobile app interfaces, multiple languages, and is adapted to country and healthcare specialty requirements.
  • Our stack is composed of Rails, TypeScript, Java, Python, Kotlin, Swift, and React Native.
  • We leverage AI ethically across our products to empower patients and health professionals. Discover our AI vision here.

Want to learn more about our tech culture and environment? Visit the Doctolib Tech site.

What we offer

  • Free comprehensive health insurance (basic package) for you and your children
  • 25 days of paid vacation per year, plus up to 14 days of RTT
  • Free mental health and coaching services through our partner Moka.care
  • Work from abroad for up to 10 days per year thanks to our flexibility days policy
  • Lunch vouchers (Swile card) worth 8.50 euros per working day, with 4.50 euros covered by Doctolib
  • 50% reimbursement of your public transport subscription
  • Parent Care Program: receive one additional month of leave on top of the legal parental leave
  • Enrollment in Doctolib's long-term employee value sharing plan called DoctoGrowth
  • Relocation support in case of international mobility
  • Access to the best AI tools for coding, development and dedicated training

Our interview process

  • Recruiter Interview
  • System Design Interview
  • Technical SRE Interview
  • Behavioral Interview
  • At least one reference check

We want your experience to be clear, respectful, and transparent. Learn more about our hiring process on our candidate experience page.

Job details

  • Permanent position
  • Tech stack: Kubernetes, Terraform, AWS / GCP, Prometheus, OpenTelemetry, Datadog, ArgoCD
  • Full-time
  • Paris, France
  • Hybrid work setup (up to 2 remote days per week)
  • Start date: as soon as possible

We welcome everyone

At Doctolib, we are committed to improving access to healthcare for everyone. This translates into our recruitment process. We evaluate candidates based solely on qualifications and motivation, without any form of discrimination.

The more diverse ideas are heard, the more our product will truly improve healthcare for all. You are welcome to apply to Doctolib, regardless of your gender, religion, age, sexual orientation, ethnicity, or disability.

To ensure equal opportunities, we invite you to exclude personal information (e.g., pictures, age) from your applications. If you require any accommodation, please let us know for support during the hiring process.

Join us in building the healthcare we all dream of!

Your data privacy

All information provided is processed by Doctolib for application management. For data processing details, click here:

Germany
|

France
|

Italy
|

Netherlands
. Please contact hr.dataprivacy(at)doctolib.com for inquiries or to exercise your rights.

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

Site Reliability Engineer

Lead SRE Network

Thales

Parislead

Откликнуться на сайте работодателя

Описание вакансии

QUI SOMMES-NOUS ?
S3NS
est né du partenariat industriel entre Thales, leader mondial de la cyber sécurité, et Google Cloud, leader mondial des solutions cloud. Nous avons pour ambition d’offrir le meilleur des deux mondes à l’ensemble des organisations soucieuses de protéger leurs données sensibles (institutions publiques, OIV, OSE…). C’est-à-dire une solution équivalente à Google Cloud Platform (incluant à la fois les services IaaS et PaaS de GCP) et respectant les exigences du label SecNumCloud.

Une première offre, ‘Contrôles locaux avec S3NS’, est déjà disponible depuis février 2023 pour permettre à nos clients de bénéficier d’un premier niveau de transparence et contrôles additionnels, et d'accélérer la trajectoire vers le cloud de confiance.

Au cœur de S3NS, le Team Lead Network occupe une position stratégique alliant leadership humain et excellence technologique. Vous ne pilotez pas seulement des infrastructures ; vous guidez une équipe de 10 experts dans la construction et l'exploitation du premier Cloud de Confiance français. Vous êtes le garant de la cohésion de l’équipe, de la fluidité des rituels et de la robustesse d’un réseau conçu pour répondre aux exigences les plus strictes de l'ANSSI (SecNumCloud). Au cœur du partenariat stratégique avec Google Cloud. Vous aurez l'opportunité unique d'explorer et de maîtriser l'architecture profonde du réseau de Google

Les missions :

Management, RH & Leadership (50%)

  • Animation d'Équipe : Manager et faire progresser une équipe de 10 ingénieurs. Garantir un environnement de travail fondé sur l'entraide et l'excellence technique.
  • Rituels & Backlog : Animer les cérémonies Agiles. Arbitrer les priorités entre les projets de construction (Build) et la stabilité opérationnelle (Run).
  • Culture SLO (Service Level Objectives) : Définir, suivre et rapporter les indicateurs de performance (SLIs/SLOs/SLAs). Mettre en place des revues périodiques pour garantir que le réseau répond aux engagements de service S3NS.
  • Gestion RH & Astreintes : Piloter la stratégie de croissance de l'équipe. Accompagner le développement professionnel de vos collaborateurs (coaching, mentors, plans de formation). Réaliser les entretiens annuels et définir les trajectoires d'évolution pour fidéliser les expertises clés. Structurer et superviser le dispositif d'astreinte pour garantir une réactivité 24/7 sur un environnement de production critique.
  • Gestion Budgétaire : Suivre et optimiser les coûts (OPEX/CAPEX) liés aux infrastructures réseau, suivi des appels d’offres et commandes.
  • Partenariat stratégique avec Google Cloud : Coordonner les activités avec les équipes réseau de Google. Aligner les besoins de sécurité et opérationnel de S3NS avec le partenaire.
  • Expertise Network & Industrialisation, Automation(30%)
  • Maîtrise des fondations GCP : Collaborer étroitement avec les équipes d'ingénierie de Google pour appréhender et opérer les couches réseau complexes (VPC, Cloud Interconnect, Andromeda, Cloud Armor).
  • Infrastructure Core : Garantir la robustesse de la couche physique et du routage : Maîtriser les réseau de Datacenter, VXLAN/EVPN, BGP haute performance et segmentation avancée. Maîtriser IPV6. Avoir une bonne culture WAN (SRV6)
  • Maîtrise des Flux : Superviser les architectures de Firewall, Forward et Reverse Proxy (WAF, Load Balancing, inspection SSL) pour assurer un contrôle total des flux entrants et sortants.
  • Infrastructure as Code (IaC) : Porter la vision "Everything as Code" pour garantir des déploiements reproductibles et auditables.
  • Stack Technologique :


Terraform
: Pour le provisionnement des infrastructures hybrides (Cloud et On-prem).


Ansible
: Pour la gestion de configuration et l'orchestration des équipements réseau.


Python
: Pour le développement d'outillages internes, d'APIs et l'automatisation de workflows complexes.


Nix / NixOS
: Exploiter Nix pour garantir la reproductibilité des environnements de build et la gestion déclarative des packages, assurant ainsi une conformité SecNumCloud inébranlable..

  • Excellence Opérationnelle, SRE & Résilience (20%)
  • Observabilité & Alerting : Concevoir une stratégie de monitoring exhaustive (Prometheus, Grafana…). Mettre en place un alerting intelligent et proactif fondé sur les indicateurs de performance clés.
  • Suivi des SLO (Service Level Objectives) : Piloter l'activité par la donnée. Garantir le respect des engagements de service (SLAs) et piloter le "budget d'erreur" pour assurer une fiabilité exemplaire.
  • Stratégie DRP/PCA : Concevoir et tester régulièrement les plans de reprise d'activité (Disaster Recovery Plan). Garantir que l'architecture réseau permet une bascule transparente ou une reconstruction rapide en cas de défaillance majeure.
  • Chaos Engineering : Initier des tests de résilience pour valider la robustesse des configurations face aux pannes d'équipements ou de liaison

Votre profil :

  • Vous êtes expert en networking (VXLAN, BGP, routage dynamique, SRv6, IPv6) et maîtrisez les interconnexions Cloud comme Interconnect ou le Peering?
  • Les architectures de Proxy (HAProxy, NGINX) et les services de sécurité périmétrique n'ont plus de secret pour vous?
  • Vous maîtrisez les outils de monitoring (Prometheus, Grafana) ainsi que les concepts de Golden Signals (Latence, Trafic, Erreurs, Saturation) , et vous savez piloter une infrastructure par la donnée et les indicateurs de fiabilité (SRE & SLOs)?
  • Vous êtes rompu aux déploiements automatisés via Terraform, Ansible, Python, les pipelines CI/CD et l'écosystème NIX?
  • Vous possédez de solides compétences en sécurité (Firewalling Next-Gen, ISE) et vous êtes familier avec les contraintes de la conformité SecNumCloud?
  • Vous disposez d'un leadership inspirant capable de fédérer une équipe autour d'une vision technologique exigeante , tout en plaçant l'humain au centre en donnant à vos collaborateurs les moyens d'exercer leur travail dans les meilleures conditions possibles?
  • Vous savez gérer des situations de crise ou des incidents majeurs sur une production critique?

Le mot de l’équipe :
"L'équipe SRE de S3NS est le moteur de la fiabilité de notre cloud de confiance. Si vous êtes passionné par la construction de systèmes robustes et performants, venez contribuer à notre mission, votre expertise fera la différence. »

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

Site Reliability Engineer

Ingénieur système RUN (H/F)

Tenexa

Colombessenior

Откликнуться на сайте работодателя

Описание вакансии

Chez Tenexa, ESN spécialisée dans l’IT, nous accompagnons nos clients sur des projets à forte valeur ajoutée en plaçant l’expertise et l’humain au cœur de nos interventions.

Dans le cadre de notre développement, nous recherchons un(e)
Ingénieur Système RUN
(H/F) pour rejoindre notre équipe à
Colombes
.

Vos missions.

Rattaché(e) à la Direction Technique et Innovation, vous intégrez l’équipe OPS en charge du RUN du Cloud Horizon.

En lien avec les équipes internes, les centres de services d’infogérance et certains clients grands comptes, vous contribuez à garantir la disponibilité, la performance et la sécurité des services Cloud.

À ce titre, vous serez amené(e) à :

  • Assurer le maintien en conditions opérationnelles du système d’information supportant les services de l’offre Cloud Horizon.
  • Traiter les demandes et incidents complexes escaladés par les équipes Front Office.
  • Instruire et mettre en œuvre les changements d’infrastructure d’origine client ou interne.
  • Piloter les analyses de causes racines et contribuer à la résolution des problèmes dans une démarche d’amélioration continue.
  • Participer aux transitions de projets dans le cadre des activités « Build to Run ».
  • Réaliser des transferts de compétences et accompagner les équipes d’exploitation sur les procédures techniques.
  • Contribuer à l’automatisation des tâches récurrentes via le scripting et les outils d’automatisation.
  • Participer aux astreintes et interventions en heures non ouvrées lorsque la criticité des services l’exige.

Vous disposez d’environ 5 ans d’expérience sur un poste d’administrateur ou d’ingénieur systèmes, idéalement au sein d’environnements cloud ou d’infrastructures complexes.

Compétences techniques requises

Systèmes

  • Windows Server (Active Directory, PKI, RDS, WSFC…)
  • Linux (Red Hat ou équivalent)
  • Administration et exploitation de serveurs physiques

Virtualisation

  • VMware ESXi / vCenter

Sauvegarde & Stockage

  • Veeam Backup & Replication
  • SAN (iSCSI, Fibre Channel)

Conteneurisation

  • Docker / Podman
  • Kubernetes apprécié

Automatisation

  • PowerShell
  • Bash ou Python
  • Ansible

Bases de données

  • SQL Server
  • MySQL / MariaDB

Gestion de parc et patch management

  • Ivanti
  • Red Hat Satellite / Uyuni

Sécurité

  • Bonnes pratiques de sécurité et référentiels ISO 27001
  • Gestion des certificats et infrastructures PKI
  • Sensibilité aux enjeux de conformité et de cybersécurité dans un contexte cloud

Compétences appréciées

  • Interventions en datacenter
  • HPE Morpheus
  • NetVault ou Commvault
  • eObserve / ServiceNav
  • Connaissances réseaux et firewall

Qualités personnelles

  • Curiosité technique et envie d’apprendre
  • Force de proposition et sens de l’amélioration continue
  • Esprit d’équipe et capacité à partager ses connaissances
  • Rigueur, organisation et esprit d’analyse
  • Bonne gestion des priorités et des situations critiques

Ce que nous proposons

  • Des missions variées et adaptées à votre profil
  • Un environnement technique stimulant autour des technologies Cloud et Infrastructure
  • Un accompagnement dans votre montée en compétences
  • Une ambiance d’équipe conviviale et un suivi régulier
  • L’opportunité de contribuer à l’évolution et à la fiabilité de services critiques

Informations complémentaires

  • Poste basé à Colombes (92)
  • Démarrage à compter de juillet 2026
  • Environ 10 % de l’activité réalisée en heures non ouvrées ou le week-end
  • Astreintes régulières à prévoir

Envie d’en savoir plus ? Postulez, et discutons ensemble de votre projet professionnel.

L'entreprise en quelques mots
Rejoindre TENEXA…
Depuis 1985, Tenexa s’est imposé comme un acteur incontournable des services numériques. Le groupe accompagne les entreprises dans leur transformation digitale en offrant des solutions sur-mesure et innovante en Cloud, Cybersécurité et Infogérance IT, répondant aux enjeux des PME et ETI d’aujourd’hui et de demain.

Une expertise à 360°
Tenexa couvre l’ensemble du cycle de vie des services IT, du DataCenter jusqu’à l’utilisateur final, garantissant performance et sécurité aux organisations de toutes tailles.

Sa stratégie repose sur une innovation continue et une excellence opérationnelle éprouvée, au service de plus de 600 clients, dont des ETI et de grands groupes issus de divers secteurs.

Une ambition soutenue par un actionnaire de référence
Le Groupe Tenexa bénéficie du soutien de HLD, société d’investissement à capitaux permanents, dédiée aux entreprises européennes à fort potentiel. Cet appui stratégique renforce la vision à long terme de Tenexa et accélère son développement sur des marchés clés.

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

Site Reliability Engineer

Site Reliability Engineer (m/f/d)

Allianz Partners

Saint-Ouenfulltime

Откликнуться на сайте работодателя

Описание вакансии

Key Responsibilities
As a Site Reliability Engineer within Advanced Analytics (DA3) in the Chief Data & AI Office at Allianz Partners, you will join the platform engineering team to own the reliability and operational health of the central engineering platform.

You will define and maintain service level objectives, drive incident response at the infrastructure layer, and systematically eliminate operational toil through automation.

You will work closely with Platform Engineers, Security Engineers, and incident-response leads to ensure the platform meets its reliability commitments across production workloads spanning AI services, Java APIs, and frontend applications.

Through this role, you will have the main following responsibilities:

  • Define, instrument, and maintain SLOs and SLIs for platform components; own error budget tracking and produce regular reliability reports for senior leadership.
  • Serve on the on-call rotation as the infrastructure escalation tier; lead incident response for cluster-level, network-level, and storage failures; chair blameless post-incident reviews.
  • Implement and operate Kubernetes infrastructure (AKS): cluster lifecycle management, networking, resource quotas, autoscaling configuration, and multi-tenancy patterns across product team namespaces.
  • Develop Infrastructure as Code (Terraform) to provision and manage Azure resources with consistency, auditability, and repeatable rollback capability.
  • Build and maintain observability infrastructure: Prometheus, Grafana, Azure Monitor, and Application Insights; own alerting rules, dashboards, and distributed tracing coverage across platform components.
  • Perform capacity planning and cost-aware resource management: right-size node pools, tune vertical and horizontal pod autoscalers, and identify resource waste across namespaces.
  • Identify and eliminate toil: automate repetitive operational tasks through scripting and tooling; measure and track toil reduction over time.
  • Maintain platform reliability procedures: rolling upgrades, backup and recovery testing, disaster recovery runbooks, and change freeze coordination.
  • Contribute to CI/CD pipelines and GitOps tooling (GitHub Actions, ArgoCD) from a reliability and deployment safety perspective; work with platform engineering on release gates and rollback mechanisms.
  • Collaborate with incident-response leads on incident SLA targets and operational procedures; work with Security Engineers on infrastructure hardening and vulnerability remediation.

What You Bring

  • 5+ years professional experience in site reliability engineering, DevOps, or platform engineering roles.
  • Strong Kubernetes experience: cluster operations, networking (Ingress, network policies), storage, autoscaling, and hands-on troubleshooting across production environments.
  • Solid Infrastructure as Code experience with Terraform; familiarity with Bicep or ARM templates is a plus.
  • Production experience with Azure cloud services: AKS, ACR, Key Vault, Azure Monitor, Application Insights, Virtual Networks, and Private Endpoints.
  • Strong observability experience: Prometheus, Grafana, centralized logging, alerting configuration, and distributed tracing instrumentation.
  • Working knowledge of SLO/SLI methodology: error budget principles, reliability target setting, and capacity planning.
  • Structured incident management experience: on-call ownership, blameless post-incident review, and runbook authorship.
  • Scripting and automation proficiency in Python or bash for toil elimination and operational tooling.
  • Strong CI/CD experience: GitHub Actions and ArgoCD or equivalent GitOps tooling.

Ways of Working

  • Comfortable in agile, iterative delivery environments with personal ownership and accountability for platform reliability.
  • Clear communicator across global, cross-functional stakeholders; able to translate technical reliability metrics into business impact for non-technical audiences.
  • Proactive learner with pragmatic adoption of AI-assisted developer tools (e.g., GitHub Copilot, Claude Code) to improve automation coverage and delivery velocity.

Nice to Have

  • Kubernetes certifications: CKA or CKAD.
  • Experience supporting AI or ML infrastructure workloads: GPU scheduling, model serving platforms, or inference pipeline operations.
  • Exposure to chaos engineering practices and fault injection testing.
  • FinOps experience: reserved capacity planning, resource right-sizing programs, and cost attribution per team or workload.
  • Service mesh experience (Istio, Linkerd) for traffic management and reliability patterns.
  • Experience in regulated industries (insurance, finance, healthcare) where auditability, change traceability, and secure-by-default operations are standard practice.

How We Hire
Allianz Partners does not accept unsolicited CV’s or approaches from agencies. We only work with partners on our approved supplier lists, under contract. Any unsolicited submission will not be considered.

What We Offer
Our employees play an integral part in our success as a business. We appreciate that each of our employees are unique and have unique needs, ambitions and we enjoy being a part of their journey. We are there to empower and encourage you with your personal and professional development ensuring that you take control by offering a large variety of courses and targeted development programs.

All that in a global environment where international mobility and career progression are encouraged. Caring for your health and wellbeing is key priority for us. This is why we build Work Well programs to providing you with peace of mind and give the flexibility in planning and arranging for a better work-life balance.

90377 | Data & AI | Professional | Allianz Partners | Full-Time | Permanent

Allianz Group is one of the most trusted insurance and asset management companies in the world. Caring for our employees, their ambitions, dreams and challenges, is what makes us a unique employer. Together we can build an environment where everyone feels empowered and has the confidence to explore, to grow and to shape a better future for our customers and the world around us.

At Allianz, we stand for unity: we believe that a united world is a more prosperous world, and we are dedicated to consistently advocating for equal opportunities for all. And the foundation for this is our inclusive workplace, where people and performance both matter, and nurtures a culture grounded in integrity, fairness, inclusion and trust.

We therefore welcome applications regardless of ethnicity or cultural background, age, gender, nationality, religion, social class, disability or sexual orientation, or any other characteristics protected under applicable local laws and regulations.

Join us. Let's care for tomorrow.

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

Ещё 54 вакансии по этой категории в этой стране

Site Reliability Engineer - StorageProton · France

Откликнуться на сайте работодателя