Skip to content
Site Reliability Engineer

Site Reliability Engineer

Site Reliability Engineer

NTT DATA Europe & Latam

Bucharestfulltime

Apply on the employer's site

Role description

Who We Are
We are looking for a hands-on
Site Reliability Engineer (SRE)
to help improve, scale, and operationalize an internal platform that enables engineering teams to ship faster and safer through reliable, automated, and resilient engineering workflows.

You will work on platform reliability, observability, automation, incident management, CI/CD workflows, GitHub-based engineering automation, and developer tooling that helps teams consistently adopt operational best practices across repositories and services.

What You'll Be Doing

  • Design, implement and maintain monitoring, alerting, and observability solutions to ensure platform reliability and performance
  • Develop and improve automation for infrastructure provisioning, deployment pipelines, and operational processes
  • Administer and optimize GitHub Enterprise environments, including repository management, access controls, branch protection policies, and enterprise-wide standards
  • Partner with engineering teams to define and measure Service Level Objectives (SLOs), Service Level Indicators (SLIs), and reliability metrics
  • Investigate production incidents, perform root cause analysis, and drive post-incident improvements to prevent recurrence
  • Improve system resilience, scalability, and availability through proactive reliability engineering practices
  • Build and maintain GitHub-based automation, CI/CD pipelines, and developer self-service capabilities

What You'll Bring Along

  • BSc/MSc in Computer Science or related field
  • Minimum 6+ years as a SRE
  • Strong experience with GitHub Enterprise (repos, orgs, actions, integrations)
  • Advanced hands-on knowledge of Terraform (IaC)
  • Experience building and running CI/CD pipelines (GitHub Actions ideally)
  • Solid understanding of cloud platforms (Azure/AWS/GCP) and integrations
  • Experience with identity and access management (RBAC, token/app auth models)
  • Knowledge of backup, recovery, and cyber resilience principles (Cybervault)
  • Experience with security best practices and software supply chain risks (due to recent events)
  • Focus on improving developer experience and self-service capabilities
  • Excellent problem-solving and root cause analysis skills

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

Senior Site Reliability Engineer

Resideo

Bucharestfulltime

Apply on the employer's site

Role description

Job Description
As a Senior Site Reliability Engineer, you will apply advanced expertise to strategically evolve and manage our cloud infrastructure, ensuring high levels of availability, scalability, and resilience. This role focuses on leading complex initiatives to drive the adoption of Infrastructure-as-Code and integrating advanced DevOps practices across global engineering teams.

Your contributions will directly influence technical architecture, operational frameworks, and the seamless delivery of cutting-edge IoT and big data platforms, requiring autonomous decision-making and cross-functional collaboration to achieve significant organizational impact.

JOB DUTIES:

  • Lead the strategic design, implementation, and optimization of public cloud infrastructure across Azure, AWS, or GCP, ensuring solutions align with organizational resilience and scalability objectives
  • Drive the adoption and continuous enhancement of Infrastructure-as-Code (IaC) principles and tools (e.g., Terraform, ARM Templates) to automate cloud resource provisioning and management
  • Develop and integrate advanced IT automation solutions using tools such as Ansible, Chef, or Helm for Kubernetes, focusing on efficiency and system reliability
  • Oversee the end-to-end CI/CD pipeline, implementing modern practices with Git, Git Actions, Jenkins, Docker, and Kubernetes to streamline software delivery and deployment
  • Architect and maintain comprehensive observability and monitoring frameworks utilizing platforms like Grafana, Prometheus, or Elastic, providing strategic insights into system performance and health
  • Manage critical incidents, conduct Root Cause Analysis (RCA) for complex outages, and lead major infrastructure upgrades to minimize downtime and ensure service continuity
  • Influence and mentor engineering teams on best practices for cloud infrastructure, reliability engineering, and operational excellence, fostering a culture of continuous improvement

YOU MUST HAVE:

  • Minimum of 6 years of progressive experience in Site Reliability Engineering or a closely related cloud infrastructure role
  • At least 3 years of hands-on experience with a major public cloud platform (Azure, AWS, or GCP), with demonstrated ability to architect and manage cloud-native solutions
  • Proven track record of designing and implementing Infrastructure-as-Code solutions for complex environments, including a minimum of 2 years with Terraform or similar IaC platforms
  • Demonstrated expertise with container orchestration platforms (Docker, Kubernetes) and supporting ecosystem

WE VALUE:

  • Advanced scripting capabilities in PowerShell, Bash, Python, or similar languages for automation and system management
  • Experience with distributed systems, large-scale data platforms, or IoT infrastructure
  • Background in a global team environment, providing strategic technical guidance and cross-functional leadership
  • Extensive experience in administering and optimizing enterprise-grade Windows and Linux environments (5+ years)
  • Strong leadership in incident response, root cause analysis, and problem resolution for complex production issues

WHAT'S IN IT FOR YOU:

  • Flexible hybrid working arrangement to support work-life balance
  • Meal ticket for each day worked
  • Medical coverage to support your health and wellbeing
  • 26 days of vacation

About Us
Resideo is a global leader in smart home and building solutions, with trusted brands including Honeywell Home, First Alert, and Resideo helping people feel more comfortable, secure, connected, and in control every day. Our products and technologies are found in more than 150 million homes and businesses worldwide. From intelligent climate solutions to security, sensing, water, and connected home technologies, Resideo develops and manufactures products designed to simplify everyday life and help protect what matters most. Our global teams span engineering, manufacturing, software, product management, supply chain, customer experience, and more — all working together to shape the future of connected living through innovation, quality, and meaningful real-world impact. At Resideo, our teams help create products and experiences that make everyday life more comfortable, secure, and connected for millions around the world. Learn more at www.resideo.com .

You can find out more about how the talent community works here: Resideo Talent Community Terms . Our recruitment privacy notice Resideo -Recruitment Privacy Notice - Dec 16 2022 describes in more detail how we process your personal data and how you can exercise your personal data rights.

If a disability prevents you from applying for a job through our website, request assistance here.

JOB INFO
Job Identification : 18846

Job Category : Software Engineering

Posting Date : 2026-06-26T14:02:52+00:00

Job Schedule : Full time

Locations : Ing. george Constantinescu 4B, Bucharest, 020339, RO

(Hybrid)

Incentive Eligible : N/A

Business : Resideo

Hiring Salary Range : At Resideo, we are committed to inclusive and equitable compensation. Salaries are determined by factors like role responsibilities, candidate qualifications, and geographic location. We also provide additional benefits tailored to your location and role.

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

Site Reliability Engineer - Observability

Adobe

Bucharestfulltime

Apply on the employer's site

Role description

The Team
We are a globally distributed team inside Adobe Developer Platforms. We own the observability platform Adobe's engineering ecosystem runs on – one of the largest logging deployments in the industry, spanning 50,000+ nodes, billions of OTel events daily, and an auth service we built in Go that handles 15,000 req/sec. Every Adobe engineer depends on what we ship.

We build it, we ship it, we run it.

What You Will Do

  • Design and operate Adobe's observability infrastructure at scale – Splunk, Loki, ClickHouse, Cortex, Tempo, and OTel collector pipelines processing billions of events daily
  • Drive cost optimisation: analyse pipeline data, build guardrails, and engage engineering teams
  • Build AI-powered tooling that surfaces actionable insights from high-volume log datasets and automates routine platform workflows
  • Own SLOs and SLIs for a platform thousands of engineers rely on every day
  • Own on-call, lead incident response for high-severity issues, and make every shift better than the last

What We Look For

  • 5–8+ years of professional DevOps or SRE experience, with a track record of owning production systems at scale
  • Deep hands-on experience with large-scale logging infrastructure – Splunk, Loki, ClickHouse, or Elastic
  • Solid working knowledge of OpenTelemetry: collector configuration, pipelines, and instrumentation
  • Strong programming skills in Go and/or Python; experience building integrations and applications to large-scale Observability environments
  • Experience developing, deploying and running distributed applications on cloud platforms; experience with container and orchestration technologies (Docker, Kubernetes)
  • Comfortable owning on-call across a multi-tool observability stack, including leading Sev 1/2 incident response

Nice to Have

  • Experience evaluating and prototyping alternative storage/processing backends (e.g., ClickHouse, Loki) as part of a cost or stability migration
  • Experience with other Observability tooling like Grafana, Cortex, and Tempo
  • Promote the DevOps/SRE approach

About Adobe
Adobe empowers everyone to create through innovative platforms and tools that unleash creativity, productivity and personalized customer experiences. Adobe’s industry-leading offerings including Adobe Acrobat Studio, Adobe Express, Adobe Firefly, Creative Cloud, Adobe Experience Platform, Adobe Experience Manager, and GenStudio enable people and businesses to turn ideas into impact, powered by AI and driven by human ingenuity.

Our 30,000+ employees worldwide are creating the future and raising the bar as we drive the next decade of growth. We’re on a mission to hire the very best and believe in creating a company culture where all employees are empowered to make an impact. At Adobe, we believe that great ideas can come from anywhere in the organization. The next big idea could be yours.

Let’s Adobe together
At Adobe, we believe in creating a company culture where all employees are empowered to make an impact. Learn more about Adobe life, including our values and culture, focus on people, purpose and community, Adobe for All, comprehensive benefits programs, the stories we tell, the customers we serve, and how you can help us advance our mission of empowering everyone to create.

Adobe is proud to be an Equal Employment Opportunity employer. We do not discriminate based on gender, race or color, ethnicity or national origin, age, disability, religion, sexual orientation, gender identity or expression, veteran status, or any other protected characteristic. Learn more.

Adobe aims to make our Careers website and recruiting process accessible to any and all users. If you have a disability or special need that requires accommodation to navigate our website or complete the application process, email accommodations@adobe.com or call +1 408-536-3015.

AI Use Guidelines for Interviews:
Our interviews are designed to reflect your own skills and thinking. The use of AI or recording tools during live interviews is not permitted unless explicitly invited by the interviewer or approved in advance as part of a reasonable accommodation. If these tools are used inappropriately or in a way that misrepresents your work, your application may not move forward in the process.

At Adobe, we empower employees to innovate with AI — and we look for candidates eager to do the same. As part of the hiring experience, we provide clear guidance on where AI is encouraged during the process and where it’s restricted during live interviews. See how we think about AI in the hiring experience.

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

Site Reliability Engineer

LSEG

Bucharestfulltime

Apply on the employer's site

Role description

ROLE SUMMARY:

We are hiring a Site Reliability Engineer to help us run and improve LSEG's internal observability platform - and we want you to be part of building something that matters!

Our platform brings telemetry, dashboards, alerts, service health, and operational insight into one common experience. We use it to help engineering and service teams detect issues earlier, reduce incident impact, and improve operational reliability.

The focus of this role is the platform itself - its reliability, supportability, and usability - and the standards and self-service patterns that help teams adopt it consistently. This is not a general application monitoring role; ownership of dashboards, alerts, or telemetry for every application team is not what we are looking for here.

WHAT YOU'LL BE DOING:

Platform operations and reliability:

  • Monitor platform health, investigate and diagnose issues, and support service recovery across production and non-production environments.
  • Contribute to service readiness through dashboards, SLOs, SLIs, runbooks, resilience checks, and performance validation.
  • Use metrics, logs, traces, and service health data to understand issues and guide practical decisions.

Incident response and continuous improvement:

  • Support incident response, triage, problem management, and root cause analysis.
  • Drive follow-up actions that reduce repeat issues and improve operational consistency.
  • Improve automation and GitOps-based workflows to reduce manual work across the team.

Enablement and documentation:

  • Maintain documentation, onboarding guides, and self-service tooling so teams can operate with less reliance on direct support.
  • Provide direct enablement on telemetry standards, alerting guidance, SLO practices, and platform workflows.

WHAT YOU'LL BRING:

Essential skills and experience:

  • Experience supporting production platforms or services in an SRE, platform engineering, infrastructure, DevOps, or operations engineering role.
  • Experience using observability data such as metrics, logs, traces, alerts, dashboards, or service health views to investigate issues.
  • Understanding of incident response, problem management, service readiness, or operational support processes.
  • Experience using automation or infrastructure-as-code practices to support repeatable delivery and operations.
  • Solid understanding of cloud, container, Linux, networking, or distributed system environments.
  • Ability to communicate technical information clearly to engineering, operations, and service stakeholders.
  • Experience writing or maintaining operational documentation such as runbooks, support guides, or onboarding material.
  • A practical approach to improving reliability, reducing manual work, and helping teams use shared platforms optimally.

Desirable skills and experience:

  • Experience with observability, monitoring, telemetry, or data pipeline technologies such as OpenTelemetry, Grafana, ClickHouse, Cribl, Datadog, BigPanda, Redis, Flink, or similar tools.
  • Experience with GitOps workflows and tools such as Git, CI/CD pipelines, pull requests, environment promotion, or configuration-as-code.
  • Experience building or supporting internal platforms used by multiple engineering teams.
  • Experience defining or using SLOs, SLIs, error budgets, alert quality measures, or service health models.
  • Experience supporting telemetry pipelines, data routing, data filtering, retention, or cost management.
  • Experience working in a regulated, financial services, or large enterprise technology environment.
  • Experience helping engineering teams adopt shared standards, templates, or self-service platform capabilities.

WHAT YOU'LL GET IN RETURN:

This is a phenomenal opportunity to work on a platform that directly improves how teams across LSEG understand and operate their systems!

We are committed to building an inclusive workplace where people from all backgrounds can grow and do meaningful work. We welcome applications from candidates with different experiences, perspectives, and career paths.

We are proud to be an equal opportunities employer. This means that we do not discriminate on the basis of anyone’s race, religion, colour, national origin, gender, sexual orientation, gender identity, gender expression, age, marital status, veteran status, pregnancy or disability, or any other basis protected under applicable law. Conforming with applicable law, we can reasonably accommodate applicants' and employees' religious practices and beliefs, as well as mental health or physical disability needs.

Career Stage:
Senior Associate

London Stock Exchange Group (LSEG) Information:
Join us and be part of a team that values innovation, quality, and continuous improvement. If you're ready to take your career to the next level and make a significant impact, we'd love to hear from you.

LSEG is a leading global financial markets infrastructure and data provider. Our purpose is driving financial stability, empowering economies and enabling customers to create sustainable growth.

Our purpose is the foundation on which our culture is built. Our values of
Integrity, Partnership
,
Excellence
and
Change
underpin our purpose and set the standard for everything we do, every day. They go to the heart of who we are and guide our decision making and everyday actions.

Working with us means that you will be part of a dynamic organisation of 25,000 people across 65 countries. However, we will value your individuality and enable you to bring your true self to work so you can help enrich our diverse workforce.

We are proud to be an equal opportunities employer. This means that we do not discriminate on the basis of anyone’s race, religion, colour, national origin, gender, sexual orientation, gender identity, gender expression, age, marital status, veteran status, pregnancy or disability, or any other basis protected under applicable law. Conforming with applicable law, we can reasonably accommodate applicants' and employees' religious practices and beliefs, as well as mental health or physical disability needs.

You will be part of a collaborative and creative culture where we encourage new ideas. We are committed to sustainability across our global business and we are proud to partner with our customers to help them meet their sustainability objectives. Our charity, the LSEG Foundation provides charitable grants to community groups that help people access economic opportunities and build a secure future with financial independence. Colleagues can get involved through fundraising and volunteering.

LSEG offers a range of tailored benefits and support, including healthcare, retirement planning, paid volunteering days and wellbeing initiatives.

Please take a moment to read this privacy notice carefully, as it describes what personal information London Stock Exchange Group (LSEG) (we) may hold about you, what it’s used for, and how it’s obtained, your rights and how to contact us as a data subject.

If you are submitting as a Recruitment Agency Partner, it is essential and your responsibility to ensure that candidates applying to LSEG are aware of this privacy notice.

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

Principal Software Engineer

Sand Technologies

Bucharestfulltime

Apply on the employer's site

Role description

Principal Software Engineer

Looking for engineering problems that actually matter?
Join Sand's Decision Intelligence Platform team and help build the semantic engines, data platforms, and AI-powered systems that enable governments and enterprises to make better operational decisions. This isn't feature development—it's designing the core technology behind critical infrastructure used by millions of people.

About Sand

Sand Technologies

is a global Physical AI company using data and AI to make critical industries work better. We partner with governments, cities and enterprises to improve how essential systems operate across healthcare, water, energy, telecommunications and infrastructure.

Our work delivers proven real-world impact. We have built AI systems that help manage London’s water supply, supported telecom network planning across hundreds of cities, and developed digital healthcare platforms serving tens of millions of people across Africa. From intelligent command centers to AI-powered infrastructure platforms, we help organizations sense, analyze and act in complex environments.

Our people are ambitious, curious and relentlessly practical. Our teams work alongside clients in the field, solving hard problems and deploying solutions that last. With colleagues across Africa, Europe, the UK and the US, we operate across the full stack - from research and engineering to deployment and capability building.

Our mission is simple: to harness AI to solve humanity’s most pressing challenges.

About the role

We're looking for a
Staff / Principal Software Engineer
to help build the next generation of Sand's
Decision Intelligence Platform
—the technology powering intelligent infrastructure for governments and enterprises across healthcare, water, telecommunications, energy, and other critical industries.

This isn't a traditional software engineering role. You'll architect and build the core capabilities that allow operators to understand complex systems, model how they behave, predict what happens next, and make better operational decisions in real time. From heterogeneous data integration and semantic modeling to distributed services, simulation engines, and AI-powered workflows, you'll own platform capabilities from concept through production.

Our platform is already solving real-world problems at national scale. Now we're looking for engineers who enjoy creating clarity from ambiguity, solving difficult systems problems, and building products where reliability matters.

What you’ll do

  • Design, build, and evolve
    core platform capabilities, including heterogeneous data integration, ontology and semantic modeling engines, simulation services, optimization capabilities, and AI-powered decision support systems.
  • Architect and develop
    scalable, production-grade distributed systems that operate across cloud-native, hybrid, sovereign cloud, and fully air-gapped environments.
  • Own the full engineering lifecycle
    for platform components—from technical design and implementation through deployment, production operations, monitoring, and continuous improvement.
  • Build robust data platforms
    capable of ingesting, transforming, and synchronizing data from modern cloud services, legacy systems, IoT devices, and disconnected operational environments.
  • Develop backend services
    using Python and Node.js/TypeScript, designing APIs and distributed services that are secure, resilient, and performant.
  • Deploy and operate AI-powered solutions
    , including machine learning and LLM-based systems, ensuring they are reliable, scalable, and production-ready.
  • Provide technical leadership
    by mentoring engineers, reviewing architecture and code, setting engineering standards, and guiding teams through complex technical decisions.
  • Collaborate closely
    with product managers, engineers, and domain experts to translate complex operational challenges into elegant technical solutions.
  • Drive engineering excellence
    by improving platform reliability, observability, automation, testing, deployment practices, and overall system performance.
  • Embrace ambiguity and ownership
    , proactively creating clarity, defining technical direction, and delivering solutions to problems that don't come with predefined answers.
  • Leverage AI throughout the engineering lifecycle
    to accelerate development, improve productivity, and build intelligent systems without compromising engineering quality.

Who you are

  • 10+ years of software engineering experience
    building, shipping, and operating production systems across multiple technology stacks and domains.
  • Strong infrastructure engineering expertise
    , with hands-on experience in Kubernetes, containers, networking, Infrastructure as Code (IaC), and production environments across AWS, Azure, and/or on-premises infrastructure.
  • Deep backend engineering experience
    using Python and/or Node.js/TypeScript, with a strong understanding of distributed systems, APIs, relational databases, and scalable application design.
  • Proven data engineering capabilities
    , including building data pipelines, streaming architectures, orchestration workflows, and integrating heterogeneous enterprise data at scale.
  • Experience deploying and operating AI/ML systems
    , including machine learning models and LLM-powered applications in production, with an understanding of MLOps principles.
  • Exceptional systems design skills
    , with a track record of architecting reliable, scalable, and resilient platforms that solve complex real-world problems.
  • A technical leader who leads by example
    , equally comfortable mentoring engineers, driving architectural decisions, reviewing code, and contributing hands-on to implementation.
  • Comfortable working across the stack
    , with enough frontend experience to design, review, and contribute when needed, while bringing deep expertise across backend and infrastructure.
  • Highly self-directed and outcome-oriented
    , able to navigate ambiguity, define technical direction, and drive complex initiatives with minimal oversight.
  • A fast learner
    who quickly adapts to new technologies, domains, and constraints, and enjoys solving problems that don't have obvious solutions.
  • An ownership mindset
    , taking responsibility for the long-term success of the systems you build, from initial design through production support and continuous improvement.
  • An AI-first engineer
    , comfortable leveraging AI-assisted development tools to accelerate delivery while maintaining high standards for software quality, architecture, and engineering excellence.

Nice to have

  • Experience building semantic data models, ontologies, or knowledge graph platforms.
  • Experience with simulation, optimization, or digital twin technologies.
  • Experience developing AI agents or autonomous workflows.
  • Exposure to healthcare, utilities, telecommunications, energy, or government technology platforms.
  • Experience building platforms for sovereign cloud or air-gapped deployments.
  • Experience contributing to platform engineering or developer tooling initiatives.

How we work

Due to the highly collaborative and internationally distributed nature of our work, successful candidates must be comfortable operating in small teams while contributing to larger, globally coordinated efforts. A strong sense of ownership, self-motivation and discipline in maintaining clear and consistent communication through virtual collaboration tools and video conferencing is essential.

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

39 more openings in this category and country

Site Reliability EngineerNTT DATA Europe & Latam · Romania

Apply on the employer's site
Site Reliability Engineer — NTT DATA Europe & Latam | mentors.coach