Skip to content
Site Reliability Engineer

Site Reliability Engineer

Especialista de SRE II

Serasa Experian

Sao Paulo

Apply on the employer's site

Role description

Company Description
A Experian é uma empresa global de dados e tecnologia que impulsiona oportunidades para pessoas e empresas ao redor do mundo. Atuamos em diversos mercados, como serviços financeiros, saúde, automotivo, agronegócio, seguros, entre outros. A Experian investe em pessoas e em novas tecnologias avançadas para liberar o poder dos dados. Contamos com uma equipe incrível de 25.200 colaboradores em 32 países.

Nossa singularidade é valorizar a sua. A cultura da Experian, centrada nas pessoas, inclusiva e orientada por propósito, é reconhecida por diversos prêmios — incluindo World’s Best Workplaces™ 2025 (Top 25 global da Fortune) e Great Place To Work™ em 26 países, entre outros. Confira o Experian Life nas redes sociais ou explore nosso site de carreiras para entender o porquê. A Experian também se orgulha de ser uma empregadora que promove igualdade de oportunidades e ação afirmativa.

Job Description
Descrição do trabalho
Estamos à procura de um Principal Site Reliability Engineer para liderar tecnicamente a disciplina de Engenharia de Confiabilidade de um produto crítico da companhia.

Esse profissional será responsável por definir a estratégia técnica de SRE, evoluir padrões de engenharia, confiabilidade, observabilidade e automação, além de atuar como referência técnica para o time e gestor das pessoas que compõem a equipe.

Esperamos um profissional com forte capacidade de liderança, visão sistêmica, excelente tomada de decisão e habilidade para influenciar diferentes áreas de tecnologia e negócio, garantindo alta disponibilidade, resiliência e eficiência operacional dos serviços.

Responsabilidades Diárias

  • Liderar técnica e gerencialmente o time de SRE, promovendo desenvolvimento técnico e de carreira dos engenheiros;
  • Definir e evoluir a estratégia de confiabilidade, observabilidade e excelência operacional dos produtos;
  • Atuar como principal referência técnica em incidentes críticos, coordenando respostas e análises de causa raiz;
  • Trabalhar em conjunto com Arquitetura, Desenvolvimento, Segurança e Produto na evolução das plataformas;
  • Conduzir rituais de operação, revisão de incidentes, capacity planning e gestão de riscos;
  • Promover cultura de automação e melhoria contínua;
  • Garantir adoção de boas práticas de SRE, DevOps e Engenharia de Plataforma;
  • Apoiar a priorização de débitos técnicos e iniciativas de confiabilidade junto às áreas de negócio.

Principais Entregas

  • Evolução dos indicadores de disponibilidade, confiabilidade e performance dos produtos;
  • Definição e acompanhamento de SLIs, SLOs e Error Budgets;
  • Evolução da plataforma de observabilidade (Logs, Métricas, Traces e APM);
  • Automação de processos operacionais reduzindo atividades manuais;
  • Redução do MTTR e aumento da capacidade de prevenção de incidentes;
  • Estruturação de Post Mortems sem culpabilização (Blameless);
  • Implantação de práticas de Chaos Engineering, Disaster Recovery e Capacity Planning;
  • Definição de padrões técnicos para infraestrutura, monitoramento e operação;
  • Desenvolvimento técnico da equipe através de mentoring, feedbacks, dojos e comunidades técnicas.

O que estamos buscando em você

Liderança

  • Experiência comprovada liderando equipes técnicas de SRE, DevOps ou Plataforma;
  • Experiência como gestor de pessoas;
  • Capacidade de influenciar diferentes áreas sem autoridade direta;
  • Excelente comunicação executiva e técnica;
  • Experiência na gestão de incidentes críticos e crises;
  • Forte capacidade analítica e tomada de decisão baseada em dados.

Qualifications
Qualificações
Conhecimentos Técnicos

  • Kubernetes (EKS/OpenShift)
  • Docker
  • AWS
  • Terraform
  • GitHub Actions, Jenkins ou outras ferramentas CI/CD
  • Observabilidade (Dynatrace, Datadog, Grafana, Prometheus, ELK, OpenTelemetry)
  • Linux
  • Redes, DNS, HTTP, TLS e balanceadores
  • Automação utilizando Python, Go ou Shell Script
  • Infraestrutura como Código (IaC)
  • Engenharia de Performance
  • Resiliência e Alta Disponibilidade
  • Gestão de Capacidade
  • Disaster Recovery
  • Chaos Engineering
  • Gestão de Vulnerabilidades
  • Práticas DevSecOps

Será um diferencial

  • Certificações AWS Professional ou Specialty;
  • Certificação Kubernetes (CKA/CKS);
  • Experiência com Service Mesh (Istio);
  • Experiência com Backstage ou Platform Engineering;
  • Conhecimento em FinOps;
  • Experiência em ambientes regulados e produtos de missão crítica;
  • Vivência com arquitetura orientada a eventos (Kafka, RabbitMQ).

Additional Information
A Serasa Experian é muito mais do que você imagina. Com o propósito de criar um futuro melhor, ampliando oportunidades para pessoas e empresas, no Brasil somos mais de 4 mil pessoas que atuam em diversos times e especialidades. Aqui, cada conhecimento e diversidade se complementa e você pode trabalhar no que mais ama, estamos comprometidos a construir uma cultura inclusiva e um ambiente no qual pessoas possam equilibrar a carreira com seus compromissos e interesses pessoais, prezando pelo bem-estar.

A gente se dedica muito em ser uma das melhores e mais inovadoras empresas para se trabalhar do país, possibilitando experiências e carreiras incríveis para nossas pessoas. Nossa forte abordagem de pessoas em primeiro lugar é reconhecida externamente por meio de diversas certificações de mercado: fomos premiados pelo Great Place To Work™ em 24 países e pela certificação internacional Top Employers, além de sermos reconhecidos como uma das melhores empresas para jovens profissionais e contarmos com uma avaliação de 4,6 no Glassdoor. Cada reconhecimento nos indica que estamos no caminho certo, proporcionando um ambiente de trabalho cada vez melhor para nossos talentos.

Experian Careers - Creating a better tomorrow together

Find out what its like to work for Experian by clicking here

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Coming to this page

A resume for this role — and a ticket to the draw

We take the posting apart down to the real requirements and rewrite your resume against it — by asking, not inventing: no line appears without your confirmation. Sign in to get it first, and to enter the draw.

  • A resume for this exact role, not a universal one
  • Answers are kept: edit any one, not the whole conversation
  • All in your account — open it from any device

On the wheel

A discount on mentoring

Winners are drawn at random among entries with a confirmed email. The date and the full rules are on the draw page.

Draw rules

Site Reliability Engineer

Desenvolvedor Fullstack (Rancher) SR

AVANTTi

Sao Paulo

Apply on the employer's site

Role description

  • Contratante: Avantti Consultoria TI
  • Atuação: 100% Remota
  • Contrato: CLT + Benefícios

Na Avantti Consultoria TI buscamos profissionais sêniores que queiram ir além do desenvolvimento convencional. Você fará parte de uma squad técnica com foco em
engenharia de software
de ponta, atuando em ambientes complexos e desafiadores, com metodologias ágeis e práticas modernas de mercado.

Nosso objetivo é garantir resiliência tecnológica, evolução contínua dos sistemas e prevenção de falhas, sempre alinhados à cultura SRE e às melhores práticas de DevSecOps.

Requisitos
🔹
Responsabilidades

  • Atuar como Desenvolvedor Full Stack em projetos técnicos.
  • Realizar análise crítica de código, propor arquitetura sistêmica e selecionar soluções modernas em ambientes TST, HML e PRD.
  • Desenvolver, refatorar e evoluir sistemas em Java (Spring Boot, Spring Batch, OpenLiberty) e Angular.
  • Garantir resolução de bugs na causa raiz.
  • Assegurar saúde e resiliência dos ambientes de produção via observabilidade, tuning de JVM e troubleshooting em clusters Kubernetes geridos pelo Rancher.
  • Atuar em prevenção/predição de falhas (SRE), gestão de capacidade (Requests/Limits), saúde de containers e expurgo otimizado de bases de dados.
  • Conduzir ações de obsolescência sistêmica, mitigação/correção de vulnerabilidades (AppSec) e otimização de imagens Docker em pipelines CI/CD.
  • Introduzir ferramentas de Inteligência Artificial (IA) para acelerar desenvolvimento, refatoração e automação de testes.

🔹
Requisitos

  • Experiência sólida em análise de código, refatoração e arquitetura sistêmica.
  • Domínio em Java (Back-End) e Angular/TypeScript (Front-End), com suporte a legado (JavaScript/jQuery/HTML).
  • Vivência com Spring Boot, Spring Batch, OpenLiberty/Jakarta EE.
  • Experiência prática com Docker, Kubernetes e Rancher (troubleshooting, monitoramento e governança).
  • Domínio em microsserviços, APIs REST e padrões de resiliência (Circuit Breakers, Retries, Graceful Shutdown, Rate Limiting).
  • Conhecimento avançado em Oracle/SQL (queries, procedures, performance, expurgo).
  • Experiência com ferramentas de CI/CD e DevSecOps: Git, GitLab, Jenkins, SonarQube, Trivy, Snyk.
  • Conhecimento prático em cultura SRE, metodologias ágeis (Scrum/Kanban) e ITIL.

🔹
Desejáveis

  • Certificações: Kubernetes CKAD/CKA, Java, Spring, Rancher, DevSecOps.
  • Conhecimento em DataStage (modelagem e engenharia de dados).
  • Experiência com Cloud (AWS, Azure, GCP).
  • Vivência com IA aplicada à engenharia de software (Copilot, geração/análise de código).

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

Support and Monitoring Analyst

IRiS Grupo Tecnológico

Sao Paulo

Apply on the employer's site

Role description

Are you passionate about cloud infrastructure and looking to be part of a dynamic team? At uCloud, we are looking for a Cloud Support & SysOps Analyst to join our global operations and monitoring team.

As part of your role, you will be responsible for:

  • Critical Monitoring & Operations: Ensure the availability and performance of cloud infrastructure by strictly following and executing documented processes (Runbooks/Playbooks).
  • Immediate Incident Response: Act with agility and dynamism in response to critical monitoring alerts and urgent infrastructure technical support requests.
  • GCP Environment Support: Provide operational support for core services such as Compute Engine, Google Kubernetes Engine (GKE), Identity and Access Management (IAM), and storage administration in Google Cloud Storage.
  • Continuous Improvement & Automation: Adopt a proactive approach to suggest improvements to the technical environment, optimize observability dashboards, and propose automations for repetitive tasks.
  • User & Workspace Support: Administer the Google Admin console and associated corporate tools to resolve incidents with a strong focus on User Experience (UX).
  • Operational Management (ITSM): Register, classify, and document the lifecycle of incidents and requests using Jira Service Manager.

Requisitos mínimos

  • Cloud Computing: Strong hands-on experience with Google Cloud Platform (GCP).
  • Cost Estimation: Ability to create budgets and projections using the Google Cloud Pricing Calculator.
  • Infrastructure & Log Analysis: Ability to perform cloud infrastructure and log analysis using native Google tools.
  • Monitoring & Observability: Consolidated experience with Zabbix and GCP tools: Google Cloud Monitoring (including implementing alert policies), configuring Uptime Checks, and advanced analysis via Logs Explorer.
  • GCP Infrastructure & Storage: Experience with VM instances (Compute Engine), GKE (Google Kubernetes Engine), permission management (IAM), database infrastructure support (Cloud SQL — without internal administration focus), and Bucket management in Google Cloud Storage (understanding classes such as Standard, Nearline, Coldline, and Archive).
  • Google Workspace Ecosystem: Operational mastery of corporate tools (Gmail, Chat, Calendar, Meet, Sheets, Docs, etc.) and experience with Google Admin to provide technical support to users, alongside daily work usage.
  • ITSM Tools: Day-to-day experience using Jira Service Manager (Atlassian Suite) or a similar tool.
  • Customer Service: Previous experience in technical support focused on User Experience (UX) and agile problem-solving.

Desirable & Differential Requirements

  • Infrastructure as Code & CLI: Desirable knowledge of Terraform and proficiency in using Cloud Shell.
  • FinOps & Cost Management: Experience in billing analysis and billing permission management, as well as consumption control through Budget Alerts configurations.
  • Cloud & Workspace Certifications:
  • Active Cloud Digital Leader or Google Cloud Certified - Associate Cloud Engineer certifications.
  • Certifications oriented toward Google Workspace administration (Major differential).
  • Data Analysis & Reporting: Ability to design operational reports and dashboards using Looker Studio (formerly Data Studio).
  • Security & Governance: Practical knowledge of secret management with Secret Manager and vulnerability monitoring through Security Command Center.
  • Serverless Architecture & Messaging: Knowledge or experience with deployments in Cloud Run and queue/event management with Pub/Sub.
  • Operating Systems: Technical knowledge in Linux environments (equivalent to LPIC-1 and LPIC-2 certification requirements).
  • Automation & Development: Basic knowledge of programming logic (in any scripting language) oriented toward automating daily processes.
  • Documentation & AI: Experience drafting internal technical documentation and using Artificial Intelligence tools to optimize routines.

Languages

  • Portuguese: Native or advanced fluency (mandatory for reading, writing, and conversation in daily internal alignments).
  • English: Intermediate reading comprehension level (focus on reading and interpreting technical IT documentation).
  • Spanish: Fluid or advanced level (completely optional/desirable, considered solely as a differential for regional communication).
  • Schedule & Shift: 12-hour alternating night shift (Standard from 18:00 to 06:00).
  • Important Note: Availability for coverage/swapping with the day shift (06:00 to 19:00) up to 2 times a week is mandatory, preferably on Thursdays and Fridays.

Benefits

  • PJ Contract
  • Health Insurance

#Ucloud

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

Back End Developer (AI Infrastructure)

Alignerr

Sao Paulo

Apply on the employer's site

Role description

About The Role
What if the code you write could directly power the AI systems shaping how millions of people interact with technology? We're looking for Back End Developers in São Paulo to design, build, and optimize the server-side systems and APIs that drive cutting-edge AI products — from data pipelines and model-serving infrastructure to scalable microservices that handle real-world traffic at scale.

This is a fully remote, flexible contract role. Whether you're a seasoned engineer looking for meaningful project work or a strong developer ready to apply your skills to AI-adjacent infrastructure, this is your chance to build things that matter.

  • Organization: Alignerr
  • Type: Hourly Contract
  • Location: Remote
  • Commitment: 10–40 hours/week

What You'll Do

  • Design, develop, and maintain robust back end services, APIs, and data pipelines
  • Build scalable, high-performance server-side architecture to support AI training and inference workloads
  • Write clean, well-documented, and testable code following engineering best practices
  • Integrate with databases, message queues, cloud services, and third-party APIs
  • Optimize system performance, reliability, and scalability as usage grows
  • Collaborate asynchronously with distributed teams of engineers, researchers, and product leads
  • Troubleshoot production issues, conduct code reviews, and contribute to architectural decisions
  • Work independently on task-based and project-based assignments on your own schedule

Who You Are

  • Proficient in one or more back end languages — Python, Go, Java, Node.js, Rust, or similar
  • Solid understanding of RESTful API design, microservices architecture, and distributed systems
  • Experienced with relational and/or NoSQL databases (PostgreSQL, MongoDB, Redis, etc.)
  • Comfortable working with cloud platforms such as AWS, GCP, or Azure
  • Familiar with version control (Git), CI/CD pipelines, and containerization (Docker, Kubernetes)
  • Strong problem-solving skills and the ability to debug complex systems independently
  • Self-motivated, reliable, and effective when working asynchronously without close supervision
  • Clear communicator who can document decisions and collaborate across time zones

Nice to Have

  • Experience building infrastructure for machine learning or AI systems
  • Familiarity with data processing frameworks (Apache Kafka, Spark, Airflow, or similar)
  • Background in performance optimization, caching strategies, or high-throughput systems
  • Knowledge of authentication, authorization, and security best practices
  • Experience with serverless architectures or event-driven design patterns
  • Contributions to open-source projects or a strong portfolio of personal/professional work

Why Join Us

  • Work on cutting-edge AI infrastructure projects alongside leading research labs
  • Fully remote and flexible — set your own hours and work from anywhere
  • Freelance autonomy with access to meaningful, technically challenging work
  • Build systems at the intersection of software engineering and artificial intelligence
  • Collaborate with a global network of talented engineers and researchers
  • Potential for ongoing work and contract extension as new projects launch

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

Gerente IT

Gallagher

Sao Paulo

Apply on the employer's site

Role description

Introduction
Welcome to Gallagher - a global community of people who bring bold ideas, deep expertise, and a shared commitment to doing what’s right. We help clients navigate complexity with confidence by empowering businesses, communities, and individuals to thrive. At Gallagher, you’ll find more than a job; you’ll find a culture built on trust, driven by collaboration, and sustained by the belief that we’re better together. Whether you join us in a client-facing role or as part of our brokerage division, our benefits and HR consulting division, or our corporate team, you’ll have the opportunity to grow your career, make an impact, and be part of something bigger. Experience a workplace where you’re encouraged to be yourself, supported to succeed, and inspired to keep learning. That’s what it means to live The Gallagher Way.

Overview
As IT Manager for Brazil, you will lead technology operations for our São Paulo and Rio de Janeiro offices. Your work will keep people connected, systems secure, and the business running smoothly. This role matters because it supports teams every day and helps the company grow with reliable, scalable technology.

You will shape how IT supports the business (today and in the future) while leading people, partners, and platforms with care and purpose

How You'll Make An Impact

  • Ensure employees have reliable, secure, and well‑supported technology to do their best work.
  • Lead and support IT analysts, helping them grow and succeed as a team.
  • Drive system integrations and technology transitions aligned with global standards.
  • Translate complex technical needs into clear, practical guidance for business leaders.
  • Improve how IT works by creating clear processes, documentation, and support models.
  • Manage technology vendors and partnerships to deliver quality and value.
  • Strengthen data protection, cybersecurity, and compliance, including Brazil’s LGPD.
  • Help modernize infrastructure through migration strategies and new technical solutions.

About You

  • Experience leading IT teams and/or technical projects in a corporate environment.
  • Professional fluency in English.
  • Strong knowledge of Microsoft‑based environments, including Windows, Microsoft 365, Azure, Active Directory, and Exchange.
  • Experience with infrastructure, networks, telephony, virtual private networks (VPNs), and system integrations.
  • Confidence supporting and resolving complex third‑level (or higher) technical issues.
  • Experience working with Quiver or similar enterprise platforms.
  • The ability to explain technical concepts clearly to both technical and non‑technical audiences.
  • Experience managing vendors, equipment inventory, and IT services.
  • The ability to prioritize multiple requests while staying organized and calm.
  • A proactive, independent working style, with good judgment on when to escalate.

Nice to have, but not required:

  • Intermediate or advanced Spanish.
  • Project Management certification.
  • Experience in insurance, reinsurance, or brokerage environments.

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

34 more openings in this category and country

Especialista de SRE IISerasa Experian · Brazil

Apply on the employer's site