Skip to content
SRE

SRE

Senior Site Reliability Engineer

TOPdesk

Budapest

Apply on the employer's site

Role description

Company Description
TOPdesk builds service management software used across education, healthcare, government, and manufacturing. We are 800+ colleagues in 11 offices worldwide. Founded over 30 years ago, we serve more than 10 million users worldwide and have been helping organisations deliver better services ever since.

We are an open, collaborative organisation with little hierarchy — people own their work end to end and are trusted to make the decisions that matter. We are reinventing ITSM and ESM for the agentic era, building AI agents our customers can trust, and this role is part of that.

Job Description
About the role

Our Azure SaaS estate keeps service management running for thousands of organisations worldwide, under SLA-backed 24/7 availability. As a Senior Site Reliability Engineer, you own the reliability of that estate as an engineering problem — you set the SLOs, engineer out the toil behind them, and make the platform faster to change and cheaper to operate without trading away resilience.

You sit in the SaaS infrastructure function, working alongside cloud engineering and the product squads shipping to production. You bring our AI-native ways of working into reliability: agents and bounded automation with observability, approvals, containment, and rollback — self-healing systems, not runbooks worked by hand.

What this is not A ticket-driven, break-fix ops role kept away from the code. This is reliability as engineering — you own SLOs and error budgets, automate what you repeat, and design the platform to recover itself rather than reacting incident by incident.

The team

We are a group of social technicians who value transparency, open feedback, and a healthy work-life balance — and who treat reliability as a shared, measurable objective, not a firefight.

What you'll own

  • SLOs and error budgets. Define and own service-level objectives across the Azure (and potentially multi-cloud) estate, and use error budgets to steer the balance between shipping change and protecting reliability.
  • Toil elimination and self-healing automation. Identify toil, classify it, and engineer it out — feeding self-healing automation and your findings into the reliability roadmap. Stand up an agent-based support layer that owns recurring toil and continuously feeds improvements back into reliability.
  • Observability consolidation. Standardise metrics, alerting, and tracing across all datacenters, close coverage gaps on cloud workloads, and measurably reduce the alert-to-incident ratio from baseline.
  • Incident response and blameless postmortems. Lead incidents to resolution, run blameless postmortems, and turn every learning into a durable fix or an automation candidate.
  • Reliability of releases. Harden CI/CD and progressive delivery — canaries, safe rollouts, automated rollback — so change velocity and reliability rise together.
  • Capacity and performance. Model capacity, load-test critical paths, and keep the platform within its performance envelope as it scales across regions.
  • AI-native reliability. Bring agents and bounded automation — with observability, approvals, containment, and rollback — into detection, diagnosis, and remediation.
  • Runbooks that get used. Every alert links to a runbook; every runbook links to an automation candidate. You leave things more legible than you found them.
  • Capacity and cost forecasting. Own capacity and cost planning across the multi-cloud estate, model usage and growth trends, and forecast short and long term infrastructure needs so spend and scaling decisions stay ahead of demand rather than reacting to it.

How you approach the work

  • Automate what you repeat — if you have done it manually twice, the third time is a design problem.
  • Measure before optimising: SLOs, baselines, and dashboards before opinions.
  • Design for failure — assume things break, and make recovery automatic and observable.
  • Consultative, not gatekeeping: you pair with product engineering teams and transfer knowledge as you go.
  • Treat cost and reliability as joint objectives, not a forced trade-off.
  • Pro-active collaboration with product teams. You are involved in the early phases of product development, including design to help the teams make optimal choices and timely introduce appropriate SRE practices.

Technical environment

  • Scale: 10+ global datacenters; SLA-backed, 24/7 multi-tenant SaaS serving millions of end users.
  • Cloud: Azure across all production regions, with a mature landing-zone and networking architecture.
  • Compute: Kubernetes / Azure AKS alongside traditional VM infrastructure, all managed as code.
  • Infrastructure as code: Terraform via CI/CD and GitOps workflows; configuration management with Puppet and Ansible across Linux and Windows.
  • Observability: metrics, alerting, and tracing across cloud-native and self-managed layers (e.g. Grafana, Prometheus, VictoriaMetrics, Influx).
  • Automation: Python and automation tooling — and we expect you to take the reliability stack to the next level, not just operate today's.
  • Legacy: Java, MS SQL, heritage architecture — being decomposed. The SRE role is not responsible for the Java application code.
  • How we build: Claude Code as our primary AI-native SDLC tool; subagents and multi-agent workflows; MCP tool integrations; shared prompt, agent, and eval libraries.

Success in your first year

  • SLOs and error budgets are defined for the estate's critical services and actively used to steer delivery decisions.
  • The alert-to-incident ratio is measurably down, and runbooks you wrote are used by the on-call shift without escalation.
  • Toil you identified is automated — or has a credible, documented roadmap to be — and self-healing covers at least one high-frequency failure mode.
  • Postmortems produce durable fixes, not repeat incidents; recurring-incident rate is trending down against a documented baseline.
  • Product squads consult you during design, not only after incidents.

Qualifications
Required

  • Proven hands-on experience (5+ years) as a Site Reliability, DevOps, or Infrastructure Engineer running a production cloud environment at scale (Azure).
  • Fluent with SLOs, error budgets, and reliability engineering practice — you have set them, not just read about them.
  • Strong observability skills at scale — Grafana, Prometheus, VictoriaMetrics, or equivalent — including alerting and tracing.
  • Kubernetes at operator level: Helm, namespace management, ingress controllers, RBAC, persistent volumes.
  • Coding for automation (Python or equivalent) and Terraform delivered via CI/CD.
  • Linux system administration — you understand what Puppet or Ansible is doing, not just whether it ran green.
  • Comfortable leading incidents in an on-call rotation with real SLA obligations, and the maturity to know when to escalate.
  • Strong written communication — your postmortems, runbooks, and architecture notes are unambiguous.

Nice to have

  • Experience with progressive delivery — canaries, feature flags, automated rollback.
  • Experience working within or migrating toward an Azure Cloud Adoption Framework or enterprise landing-zone structure.
  • Current, personal practice of AI-native software delivery (Claude Code or equivalent).
  • Experience with EU data residency / sovereign cloud requirements.

Additional Information
What's in it for you

  • Extra 5 additional days of leave on top of the basic annual leave
  • Holiday allowance: 8% of the annual gross salary (paid in December)
  • Fitness membership (Life1 gym)
  • Private health insurance (Medicare)
  • Cafeteria
  • Company massage
  • Flexible working hours
  • Home office
  • Personal development (no fixed budget)
  • Anniversary bonus
  • Sabbatical

Want to apply?
Does this sound like your next step? Apply with your CV and a short motivation letter via the application form. Tell us about a service you made more reliable — the SLOs you set, the toil you engineered out, and the outcome (uptime, incidents reduced, or recovery time improved).

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

SRE

Platform Engineer

zenga.hu

Budapest

Apply on the employer's site

Role description

Együtt vagyunk hatással!

A Zenga.hu-t működtető OTP Otthonmegoldások Kft. egy dinamikusan fejlődő csapat, amely nagyvállalati háttérrel és startup szemlélettel dolgozik azon, hogy Közép- és Kelet-Európában egyedülálló, innovatív otthonmegoldási szolgáltatásportfóliót hozzon létre. Ha szívesen dolgoznál izgalmas, valódi hatással bíró projekteken, miközben folyamatos fejlődési lehetőségek is várnak, csatlakozz hozzánk az alábbi pozícióba:
Platform Engineer
.

Ebben a szerepkörben a zenga.hu Azure-alapú ingatlanportál platform- és infrastruktúra-oldalának fejlesztéséért és üzemeltetéséért felelsz. Bár nem alkalmazásfejlesztőként dolgozol, érted az alkalmazások működését, és könnyedén megtalálod a közös szakmai hangot a fejlesztőkkel és a tesztelőkkel. Fontos célod, hogy a platform ne akadályozza, hanem kifejezetten támogassa és felgyorsítsa a fejlesztési és tesztelési folyamatokat.

Ezek a feladatok várnak nálunk:

  • Fejlesztői és tesztkörnyezetek (DEV, UAT) kialakítása, üzemeltetése és folyamatos rendelkezésre állásuk biztosítása
  • Cloud környezetek kezelése és konszolidálása
  • CI/CD pipeline-ok kialakítása és karbantartása Azure DevOps eszközökkel, valamint release-, A/B tesztelési és canary build folyamatok automatizálása
  • Kubernetes / AKS klaszterek konfigurálása, karbantartása és kapacitástervezése
  • Infrastructure as Code megoldások fejlesztése és fenntartása Terraform használatával
  • Monitoring és alerting rendszerek (Azure Monitor, Application Insights) fejlesztése és üzemeltetése
  • Önkiszolgáló eszközök biztosítása a fejlesztői csapatok számára (pl. logok, környezeti információk, self-service deployment)
  • Az infrastruktúra egyszerűsítése, átláthatóbbá tétele és költséghatékony működtetése
  • Performancia tesztelési infrastruktúra kialakítása és karbantartása
  • Szoros együttműködés a fejlesztő- és tesztelőcsapatokkal üzemeltethetőségi és architektúra kérdésekben
  • Új külső (3rd party) szolgáltatások bevezetésének szakmai előkészítése (költség, integrálhatóság, biztonság, fenntarthatóság szempontjából)
  • Közreműködés penetrációs tesztelési projektek koordinálásában

Téged várunk, ha van:

  • Legalább 3–5 év DevOps vagy platform engineering tapasztalatod
  • Magabiztos Kubernetes ismereted (AKS, Helm, resource management, troubleshooting)
  • Gyakorlatod Docker és konténer-alapú megoldások alkalmazásában
  • Tapasztalatod Azure PaaS és IaaS szolgáltatásokkal (pl. AKS, ACR, Key Vault, VNet, Azure Monitor, Application Insights)
  • Tapasztalatod Azure DevOps pipeline-ok fejlesztésében
  • Rendszerszemléletű, platformszintű gondolkodásmódod
  • Költségtudatos szemléleted, valamint képességed infrastruktúra-optimalizációs döntések meghozatalára
  • Fejlesztői empátiád: érted a fejlesztők és tesztelők igényeit, és proaktívan segíted a munkájukat
  • Önálló munkavégzési képességed, kezdeményezőkészséged, és együttműködő hozzáállásod

Előnnyel indulsz nálunk, ha van:

  • Tapasztalatod Terraform használatában (modulok, state management, remote backend)
  • Azure minősítésed (pl. AZ-104, AZ-400 vagy magasabb szint)
  • AWS ismereted (a platform részben AWS-en is fut)
  • Gyakorlatod GitOps szemléletben és eszközökben (pl. ArgoCD, Flux)
  • Tapasztalatod service mesh megoldásokkal (pl. Istio, Linkerd)
  • Kafka infrastruktúra üzemeltetésében szerzett gyakorlatod
  • Elasticsearch klaszterek üzemeltetési tapasztalata
  • FinOps ismereted, illetve tapasztalatod Azure cost management eszközökkel
  • Tapasztalatod PostgreSQL magas rendelkezésre állású konfigurációival (Azure Database for PostgreSQL Flexible Server)
  • Agilis/Scrum módszertan ismereted, valamint fejlesztői csapatokkal való szoros együttműködésben szerzett tapasztalatod

Ezért érdemes hozzánk csatlakozni:

  • Versenyképes juttatási csomag és cafeteria
  • Magán egészségbiztosítási csomag
  • Támogatás egészségpénztári és önkéntes nyugdíjcélú megtakarításhoz
  • AYCM sportpass lehetőség
  • Számos munkatársi kedvezmény
  • Hibrid munkavégzés és home office lehetőség
  • Céges laptop és mobiltelefon biztosítása
  • Stabil vállalati háttér, innovatív, startup szemléletű környezet
  • Dinamikus, modern („smart working”) közösség
  • Magas szintű szakmai tudással rendelkező, támogató és befogadó csapat
  • Sokszínű, izgalmas feladatok és korszerű technológiák
  • Felelősségteljes, önálló munkavégzés lehetősége
  • Folyamatos szakmai fejlődés és képzések támogatása
  • Udemy Business hozzáférés és szakmai konferenciákon való részvétel
  • Hosszú távú munkalehetőség és kiszámítható karrierút

Munkavégzés helye:
Budapest, XIII. kerület

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

SRE

Site Reliability Engineering Specialist

BT Group

Budapest

Apply on the employer's site

Role description

Job Title: Site Reliability Engineering Specialist

Req ID: 61169

Job Function: Software Engineering

Posting Start Date: 29/07/2026

Posting End Date:

Division: BT International

Job Location: HUN Budapest ONE

Advertised Salary: Competitive

About BT
BT Group is the UK’s leading communications group and the holding company behind some of the country’s most recognised brands – including BT, EE, Openreach and Plusnet. Our purpose is as simple as it is ambitious: we connect for good. Our customers include consumers, small, medium and large businesses, public sector organisations and other communications providers.

BT Group’s role is about setting direction, unlocking value and creating the conditions for our brands and businesses to thrive. Having come through the most capital-intensive phase of our fibre investment, our focus now is on what comes next – simplifying how we operate, using technology and AI to work smarter, and organising ourselves to serve customers better and grow sustainably. Group teams shape strategy, policy, brand, capital allocation and transformation, helping the whole organisation perform at its best.

We have a singular culture that unites all our people: we are customer-first challengers, who are committed, clear and connected. These behaviours unite us as one team to deliver for our colleagues, our customers, our stakeholders and the country. Joining BT Group means working at the heart of a business that matters to the UK, with the opportunity to shape decisions, influence outcomes and help set the future course of one of the country’s most important companies.

About The Role
As a Site Reliability Engineer (SRE) within the Network Operations team, BTI International, you will be responsible for ensuring the reliability, resilience and performance of our Global Platforms including Global Fabric. You will collaborate closely with Engineering, Product and Service teams to embed SRE principles such as automation, observability and proactive incident reduction into day to day operations. By improving how we monitor, maintain and evolve our services, you will help reduce risk, improve service quality and increase operational efficiency. Through this role, you will support BTI International’s strategy by enabling stable, secure and scalable platforms that support business growth, accelerate delivery of new capabilities, and protect customer experience.

What You Will Be Doing (Role Accountabilities)

  • Own the operational reliability, performance and resilience of the Global Fabric NaaS platform.
  • Support and troubleshoot microservices, APIs and integrations across the NaaS ecosystem.
  • Diagnose and resolve production issues across Kubernetes-hosted applications, Linux systems, networking, Kafka, APIs and service integrations.
  • Support safe, automated change into production using CI/CD, GitOps, and automated testing.
  • Improve observability, monitoring and traceability across the platform using Dynatrace, Prometheus, Grafana, Elasticsearch and Kafka.
  • Support BT’s move towards end-to-end tracing and service traceability, helping implement and improve synthetic monitoring, tracing and service flow visibility.
  • Participate in major incident resolution, root cause analysis and post-incident improvement activities.
  • Manage incidents, problems and changes through ServiceNow and track defects and improvements in Jira.
  • Drive automation through Ansible, Python, Bash or similar tooling to reduce manual effort and operational risk.
  • Mentor and support L2 engineers by improving troubleshooting practices, runbooks and operational readiness.
  • Build strong knowledge of the end-to-end customer journey and ensure operational decisions are aligned to customer impact.

What You’ll Need To Succeed (Skills & Experience)
Must have:

  • Strong Linux and system administration experience, including server and compute management.
  • Experience deploying, supporting and troubleshooting containerised applications in Kubernetes.
  • Experience using monitoring tools such as Dynatrace, Prometheus, Grafana, Elasticsearch and Kafka.
  • Experience supporting large-scale, high-availability services in an ISP, telecom, NaaS or network-centric environment.
  • Experience with CI/CD, GitOps and safe production deployments.
  • Experience with scripting and automation using Python, Bash, Ansible or similar.
  • Growth Mindset: Self-driven attitude towards learning new skills and aiding the development of others.

Desired:

  • In-depth knowledge of network protocols, including BGP, IS-IS and MPLS.
  • Understanding of synthetic monitoring, telemetry and end-to-end service visibility.
  • Experience of resilience, disaster recovery, chaos engineering or high availability testing.
  • Ability to manage incidents through ServiceNow, track defects and continuous improvements in Jira.

BT Group’s Behaviours
Customer First: Prioritize customer needs in every decision and action.

Challengers: Challenge the status quo and bring innovative ideas to life.Committed: Own outcomes and deliver with integrity.

Clear: Communicate openly and simply, ensuring alignment.

Connected: Collaborate across teams to achieve shared goals.

At BT International, our purpose is to keep the world connected. As part of BT, we build on almost 180 years of innovation and expertise to deliver secure connectivity and digital services to some of the world’s leading multinational businesses and organisations. Our customers trust us to safeguard their data, drive their digital transformation and keep their businesses running. With colleagues on the ground across the world and supporting customers wherever they need to operate, BT International offers a truly global experience. Whether it’s about providing cloud connectivity, helping organisations collaborate, or enabling innovation in cybersecurity and digital services, you’ll be part of a team that shapes how businesses succeed in a world that is being transformed by AI. If you have the drive and ambition to make an impact on a global stage, BT International is where it happens.

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

SRE

Application Onboarding Specialist

Lufthansa Systems Hungária

Budapest

Apply on the employer's site

Role description

Fasten your seatbelt! Are you ready to take on the challenge of working on a complex, large-scale system in an international team?

We are looking for colleagues to support the migration of Lufthansa’s Governor Identity Management (IDM) platform to a new target IDM solution a highly complex transformation initiative.

Your responsibilities will include

  • Supporting and managing the onboarding of applications into the IDM.
  • Monitoring the health, performance, and availability of the Identity Management platform.
  • Managing and resolving incidents, service requests and operational issues.
  • Executing operational procedures, maintenance tasks, and routine checks.
  • Performing first- and second-level troubleshooting activities.
  • Contributing to problem management and root cause analysis activities.
  • Maintaining operational documentation and runbooks.

We are looking for you if you have

  • Experience in application operations or production support environments.
  • Skills to break down complex concepts into simple, understandable language.
  • Familiarity with incident, problem, and change management processes (e.g. ServiceNow).
  • Experience supporting IAM/IDM platforms is highly desirable.
  • Experience with Garancy / Omada / One Identity / SailPoint is advantageous.
  • Familiarity with RBAC and authorization concepts.
  • Strong analytical and troubleshooting skills.
  • Ability to work in a structured operational environment.
  • Ability to answer and understand
    German
    .
  • Excellent written and spoken
    English
    skills.

What we are proud to provide you with

  • Hybrid working with even only weekly 2 office days and flexible working hours based on core time.
  • Financial benefits such as; an outstanding amount of cafeteria, referral reward and team or individual reward based on your accomplishments.
  • Corporate health insurance package at one of the largest private healthcare providers with nationwide coverage.
  • Carefully designed career paths, with the appropriate training courses and projects to accomplish your aims such as language courses, tech talks and trainings for your needs.
  • Lufthansa Group flight ticket discounts.
  • What’s more, if you are interested in our additional benefits: check out our benefits page.

Our community is characterized by mutual respect and a spirit of helpfulness. This high level of professionalism, combined with diverse personalities, will help you go through your daily life in a good mood while making the most of your skills and talents.

In addition to an inclusive and helpful atmosphere, trust is also an important value for us. You can rely on us with anything right from your first day, and we also trust you to the utmost extent.

Are you ready to fly with us? Apply now!

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

SRE

Incident Management Reliability Engineer

Sanofi

Budapest

Apply on the employer's site

Role description

Job title: Incident Management Reliability Engineer

  • Location: Budapest
  • Hybrid working

Our Team
Service Quality cultivates a culture of service excellence where quality is more than a benchmark – it's a shared purpose. Through synergistic collaboration, advanced monitoring, and empathetic customer advocacy, we strive to elevate every interaction and transform challenges into opportunities for growth.

Main Responsibilities
The Incident Management Reliability Engineer is responsible for ensuring the stability, resilience, and reliability of critical IT services. This role combines strong incident management expertise with reliability engineering principles to minimize disruptions, drive rapid recovery from major incidents, and continuously improve system performance and availability.

Incident Management

  • Lead the end-to-end management of Major Incidents (P1/P2), ensuring timely resolution and effective stakeholder communication.
  • Act as command centre lead during critical outages, coordinating across technical and business teams.
  • Ensure accurate and detailed incident documentation, including root cause, timeline and resolution steps.
  • Drive post-incident-reviews and ensure action items are implemented to prevent recurrence.
  • Maintain consistent communication and escalation processes aligned with ITSM best practices (e.g. ITIL)

Reliability Engineering

  • Collaborate with service owners and platform teams to enhance service reliability, observability, and fault tolerance.
  • Implement proactive monitoring, alerting, and automated recovery mechanisms.
  • Analyse incident trends and develop reliability improvement plans.
  • Participate in capacity planning, change reviews, and failure mode analysis to anticipate and mitigate risks.
  • Develop and track SLOs/SLIs/SLAs to measure service health and performance.

Continuous Improvemen
t

  • Partner with problem management to identify recurring issues and lead root cause elimination initiatives.
  • Automate operational tasks and enhance service recovery using scripts, runbooks, and AIOps tools.
  • Contribute to the evolution of the Major Incident Process, ensuring best practices are embedded across the organization.

Key Performance Indicators

  • Mean Time to Resolve (MTTR) and Mean Time to Detect (MTTD).
  • Reduction in number and impact of recurring incidents.
  • Adherence to SLA/SLO targets.
  • Completion rate of post-incident actions.
  • Stakeholder satisfaction and transparency during incidents.

Experience
About you

  • 8+ years' experience.

Preferred Certifications

  • ITIL v4 or Service Operations certification.
  • SRE Foundation / Practitioner certification.
  • Cloud certifications (AWS, Azure, or GCP).
  • Incident Command System (ICS) or equivalent leadership training in crisis response.

Soft Skills

  • Communication (verbal and written).

Essential Technical Skills

  • Networking

Technical Skills (nice To Have)

  • Virtualization
  • Cloud Technologies
  • Database
  • Containerization
  • Automation
  • Middleware/Scheduling
  • Infrastructure as code

Languages

  • English

Why choose us?

  • An international work environment, in which you can develop your talent and realize ideas and innovations within a competent team
  • Bring the miracles of science to life alongside a supportive, future-focused team
  • An environment based on last technologies and frequent training to reinforce your profile
  • Be part of a simpler, digital- and AI-powered business that’s rethinking how we work and engage with the world

Join an Award-Winning Team at Sanofi Budapest Hub!
Be part of something exceptional. Our Budapest Hub has been recognized for workplace excellence, innovation, and our commitment to putting people first. See the full list of our awards at the end of this posting.
Office of the Year 2025 – Evolution Award
Our most recent award – Sanofi is recognized for creating an innovative workspace that supports collaboration and adaptability.
Marketing Diamond Award 2026 – Employer Branding
One of the highest honours at Hungary's prestigious Marketing Diamond Awards, recognizing excellence in employer branding.
BSC Investor of the Year 2025
Awarded by HIPA, ABSL Hungary, and AmCham Hungary for our contribution to Hungary's business services sector.
PwC Workforce Preference Survey 2025 – Top 3 Most Attractive Employers
Ranked 3rd among the most attractive employers in Hungary.
Hungarian Employer Branding Awards 2025 – Gold & Silver
Best Strategic Employer Branding Campaign (Gold) and Best Employer Branding Campaign in the Pharma Sector (Silver).
#Sanofi #WeNeverSettle #SanofiCareers #PursueProgress #DiscoverExtraordinary #joinsanofi #careerswithpurpose #SBSBUDAPEST

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

14 more openings in this category and country

Senior Site Reliability EngineerTOPdesk · Hungary

Apply on the employer's site
Senior Site Reliability Engineer — TOPdesk | mentors.coach