Skip to content
Site Reliability Engineer

Site Reliability Engineer

Site Reliability Engineer

Thales

Sao Paulo

Apply on the employer's site

Role description

Thales people architect identity management and data protection solutions at the heart of digital security. Business and governments rely on us to bring trust to the billions of digital interactions they have with people. Our technologies and services help banks exchange funds, people cross borders, energy become smarter and much more. More than 30,000 organizations already rely on us to verify the identities of people and things, grant access to digital services, analyze vast quantities of information and encrypt data to make the connected world more secure.

Site Reliability Engineer
This position is
hybrid
model in our
Berrini
unit.

Position Summary
The candidate will be working as a SRE member who will help the organization to constantly ensure reliability, availability and performance of large-scale ODC services. SRE will work closely with development teams to design, build, and maintain scalable and reliable infrastructure, automate processes, monitor system health, and respond to incidents effectively with a mindset of efficiency on day-to-day activities. SRE will constantly adopt ITIL and Agile methodologies/processes, coaching and mentoring on best practices. Will endorse whole lifecycle over Public Cloud ensuring to meet external customer SLA and internal OLAs.

Key Areas of Responsibility

  • Develop and maintain Infrastructure as a Code and automation tools.
  • Responsible to Integrate, Operate and Support 7x24 mission critical services with 5x9 availability on public cloud.
  • Responsible to ensure tier 1 / Platinum SLAs.
  • Responsible to review technical products and understand customer requirements.
  • Responsible to perform regular tuning.
  • Able to work with distributed teams worldwide.
  • Responsible for defining business continuity strategy for Operated services over public cloud.
  • Motivate the team on daily basis through Agile ceremonies (Daily, refinement, planning...)
  • Responsible for suggesting indicators on team monitoring.
  • Responsible for facilitating exchanges with the many stakeholders.
  • Continuously improve service reliability, performance, and security of the services
  • Collaborate with Service Delivery Managers on traffic trends, analyses the impact of mid-term business changes on capacity requirements.
  • Participate in capacity management processes and security audits.
  • Design and implement changes into the systems.
  • Adapt solution parameters to make architecture evolutions.
  • Maintain and enhance internal tools to improve service industrialization.
  • Participate in the Presales, deployment and integration of the solutions form the Support perspectives.
  • Definition of production requirements.

Minimum Qualifications

  • Bachelor’s degree in information technology or a related field.
  • Experience in design, development and implementation of applications.
  • Experience in Public Cloud (GCP)
  • Minimum C1 English (Advanced Level)
  • Strong experience on Kubernetes (certification).
  • Strong experience on Apache Http Server
  • Strong experience on TLS >= 1.2
  • Ability to work SRE engineers during integration and operation project phases.
  • Strong experience working in Agile teams.
  • Experience in embedding agile performance metrics to drive accountability.
  • Effective verbal and written communication skills
  • Strong working experience on one of the scripting languages – SHELL/Python is required.
  • Experience on GCP Terraform

If you’re excited about working with Thales, but not meeting the requirements for this position, we encourage you to join our Talent Community!

What We Offer
Thales provides an extensive benefits program for all full-time employees working 30 or more hours per week and their eligible dependents, including the following:

  • Elective Health and Dental plans.
  • Retirement Savings Plan with a company contribution and a match.
  • Company paid holidays, vacation days, and paid sick leave.
  • Company provided Life Insurance.

Why Join Us?
Say HI and learn more about working at Thales
click here
.
At Thales we provide CAREERS and not only jobs. With Thales employing 80,000 employees in 68 countries our mobility policy enables thousands of employees each year to develop their careers at home and abroad, in their existing areas of expertise or by branching out into new fields. Together we believe that embracing flexibility is a smarter way of working. Great journeys start here, apply now!

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Coming to this page

A resume for this role — and a ticket to the draw

We take the posting apart down to the real requirements and rewrite your resume against it — by asking, not inventing: no line appears without your confirmation. Sign in to get it first, and to enter the draw.

  • A resume for this exact role, not a universal one
  • Answers are kept: edit any one, not the whole conversation
  • All in your account — open it from any device

On the wheel

A discount on mentoring

Winners are drawn at random among entries with a confirmed email. The date and the full rules are on the draw page.

Draw rules

Site Reliability Engineer

Desenvolvedor Fullstack (Rancher) SR

AVANTTi

Sao Paulo

Apply on the employer's site

Role description

  • Contratante: Avantti Consultoria TI
  • Atuação: 100% Remota
  • Contrato: CLT + Benefícios

Na Avantti Consultoria TI buscamos profissionais sêniores que queiram ir além do desenvolvimento convencional. Você fará parte de uma squad técnica com foco em
engenharia de software
de ponta, atuando em ambientes complexos e desafiadores, com metodologias ágeis e práticas modernas de mercado.

Nosso objetivo é garantir resiliência tecnológica, evolução contínua dos sistemas e prevenção de falhas, sempre alinhados à cultura SRE e às melhores práticas de DevSecOps.

Requisitos
🔹
Responsabilidades

  • Atuar como Desenvolvedor Full Stack em projetos técnicos.
  • Realizar análise crítica de código, propor arquitetura sistêmica e selecionar soluções modernas em ambientes TST, HML e PRD.
  • Desenvolver, refatorar e evoluir sistemas em Java (Spring Boot, Spring Batch, OpenLiberty) e Angular.
  • Garantir resolução de bugs na causa raiz.
  • Assegurar saúde e resiliência dos ambientes de produção via observabilidade, tuning de JVM e troubleshooting em clusters Kubernetes geridos pelo Rancher.
  • Atuar em prevenção/predição de falhas (SRE), gestão de capacidade (Requests/Limits), saúde de containers e expurgo otimizado de bases de dados.
  • Conduzir ações de obsolescência sistêmica, mitigação/correção de vulnerabilidades (AppSec) e otimização de imagens Docker em pipelines CI/CD.
  • Introduzir ferramentas de Inteligência Artificial (IA) para acelerar desenvolvimento, refatoração e automação de testes.

🔹
Requisitos

  • Experiência sólida em análise de código, refatoração e arquitetura sistêmica.
  • Domínio em Java (Back-End) e Angular/TypeScript (Front-End), com suporte a legado (JavaScript/jQuery/HTML).
  • Vivência com Spring Boot, Spring Batch, OpenLiberty/Jakarta EE.
  • Experiência prática com Docker, Kubernetes e Rancher (troubleshooting, monitoramento e governança).
  • Domínio em microsserviços, APIs REST e padrões de resiliência (Circuit Breakers, Retries, Graceful Shutdown, Rate Limiting).
  • Conhecimento avançado em Oracle/SQL (queries, procedures, performance, expurgo).
  • Experiência com ferramentas de CI/CD e DevSecOps: Git, GitLab, Jenkins, SonarQube, Trivy, Snyk.
  • Conhecimento prático em cultura SRE, metodologias ágeis (Scrum/Kanban) e ITIL.

🔹
Desejáveis

  • Certificações: Kubernetes CKAD/CKA, Java, Spring, Rancher, DevSecOps.
  • Conhecimento em DataStage (modelagem e engenharia de dados).
  • Experiência com Cloud (AWS, Azure, GCP).
  • Vivência com IA aplicada à engenharia de software (Copilot, geração/análise de código).

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

Support and Monitoring Analyst

IRiS Grupo Tecnológico

Sao Paulo

Apply on the employer's site

Role description

Are you passionate about cloud infrastructure and looking to be part of a dynamic team? At uCloud, we are looking for a Cloud Support & SysOps Analyst to join our global operations and monitoring team.

As part of your role, you will be responsible for:

  • Critical Monitoring & Operations: Ensure the availability and performance of cloud infrastructure by strictly following and executing documented processes (Runbooks/Playbooks).
  • Immediate Incident Response: Act with agility and dynamism in response to critical monitoring alerts and urgent infrastructure technical support requests.
  • GCP Environment Support: Provide operational support for core services such as Compute Engine, Google Kubernetes Engine (GKE), Identity and Access Management (IAM), and storage administration in Google Cloud Storage.
  • Continuous Improvement & Automation: Adopt a proactive approach to suggest improvements to the technical environment, optimize observability dashboards, and propose automations for repetitive tasks.
  • User & Workspace Support: Administer the Google Admin console and associated corporate tools to resolve incidents with a strong focus on User Experience (UX).
  • Operational Management (ITSM): Register, classify, and document the lifecycle of incidents and requests using Jira Service Manager.

Requisitos mínimos

  • Cloud Computing: Strong hands-on experience with Google Cloud Platform (GCP).
  • Cost Estimation: Ability to create budgets and projections using the Google Cloud Pricing Calculator.
  • Infrastructure & Log Analysis: Ability to perform cloud infrastructure and log analysis using native Google tools.
  • Monitoring & Observability: Consolidated experience with Zabbix and GCP tools: Google Cloud Monitoring (including implementing alert policies), configuring Uptime Checks, and advanced analysis via Logs Explorer.
  • GCP Infrastructure & Storage: Experience with VM instances (Compute Engine), GKE (Google Kubernetes Engine), permission management (IAM), database infrastructure support (Cloud SQL — without internal administration focus), and Bucket management in Google Cloud Storage (understanding classes such as Standard, Nearline, Coldline, and Archive).
  • Google Workspace Ecosystem: Operational mastery of corporate tools (Gmail, Chat, Calendar, Meet, Sheets, Docs, etc.) and experience with Google Admin to provide technical support to users, alongside daily work usage.
  • ITSM Tools: Day-to-day experience using Jira Service Manager (Atlassian Suite) or a similar tool.
  • Customer Service: Previous experience in technical support focused on User Experience (UX) and agile problem-solving.

Desirable & Differential Requirements

  • Infrastructure as Code & CLI: Desirable knowledge of Terraform and proficiency in using Cloud Shell.
  • FinOps & Cost Management: Experience in billing analysis and billing permission management, as well as consumption control through Budget Alerts configurations.
  • Cloud & Workspace Certifications:
  • Active Cloud Digital Leader or Google Cloud Certified - Associate Cloud Engineer certifications.
  • Certifications oriented toward Google Workspace administration (Major differential).
  • Data Analysis & Reporting: Ability to design operational reports and dashboards using Looker Studio (formerly Data Studio).
  • Security & Governance: Practical knowledge of secret management with Secret Manager and vulnerability monitoring through Security Command Center.
  • Serverless Architecture & Messaging: Knowledge or experience with deployments in Cloud Run and queue/event management with Pub/Sub.
  • Operating Systems: Technical knowledge in Linux environments (equivalent to LPIC-1 and LPIC-2 certification requirements).
  • Automation & Development: Basic knowledge of programming logic (in any scripting language) oriented toward automating daily processes.
  • Documentation & AI: Experience drafting internal technical documentation and using Artificial Intelligence tools to optimize routines.

Languages

  • Portuguese: Native or advanced fluency (mandatory for reading, writing, and conversation in daily internal alignments).
  • English: Intermediate reading comprehension level (focus on reading and interpreting technical IT documentation).
  • Spanish: Fluid or advanced level (completely optional/desirable, considered solely as a differential for regional communication).
  • Schedule & Shift: 12-hour alternating night shift (Standard from 18:00 to 06:00).
  • Important Note: Availability for coverage/swapping with the day shift (06:00 to 19:00) up to 2 times a week is mandatory, preferably on Thursdays and Fridays.

Benefits

  • PJ Contract
  • Health Insurance

#Ucloud

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

Back End Developer (AI Infrastructure)

Alignerr

Sao Paulo

Apply on the employer's site

Role description

About The Role
What if the code you write could directly power the AI systems shaping how millions of people interact with technology? We're looking for Back End Developers in São Paulo to design, build, and optimize the server-side systems and APIs that drive cutting-edge AI products — from data pipelines and model-serving infrastructure to scalable microservices that handle real-world traffic at scale.

This is a fully remote, flexible contract role. Whether you're a seasoned engineer looking for meaningful project work or a strong developer ready to apply your skills to AI-adjacent infrastructure, this is your chance to build things that matter.

  • Organization: Alignerr
  • Type: Hourly Contract
  • Location: Remote
  • Commitment: 10–40 hours/week

What You'll Do

  • Design, develop, and maintain robust back end services, APIs, and data pipelines
  • Build scalable, high-performance server-side architecture to support AI training and inference workloads
  • Write clean, well-documented, and testable code following engineering best practices
  • Integrate with databases, message queues, cloud services, and third-party APIs
  • Optimize system performance, reliability, and scalability as usage grows
  • Collaborate asynchronously with distributed teams of engineers, researchers, and product leads
  • Troubleshoot production issues, conduct code reviews, and contribute to architectural decisions
  • Work independently on task-based and project-based assignments on your own schedule

Who You Are

  • Proficient in one or more back end languages — Python, Go, Java, Node.js, Rust, or similar
  • Solid understanding of RESTful API design, microservices architecture, and distributed systems
  • Experienced with relational and/or NoSQL databases (PostgreSQL, MongoDB, Redis, etc.)
  • Comfortable working with cloud platforms such as AWS, GCP, or Azure
  • Familiar with version control (Git), CI/CD pipelines, and containerization (Docker, Kubernetes)
  • Strong problem-solving skills and the ability to debug complex systems independently
  • Self-motivated, reliable, and effective when working asynchronously without close supervision
  • Clear communicator who can document decisions and collaborate across time zones

Nice to Have

  • Experience building infrastructure for machine learning or AI systems
  • Familiarity with data processing frameworks (Apache Kafka, Spark, Airflow, or similar)
  • Background in performance optimization, caching strategies, or high-throughput systems
  • Knowledge of authentication, authorization, and security best practices
  • Experience with serverless architectures or event-driven design patterns
  • Contributions to open-source projects or a strong portfolio of personal/professional work

Why Join Us

  • Work on cutting-edge AI infrastructure projects alongside leading research labs
  • Fully remote and flexible — set your own hours and work from anywhere
  • Freelance autonomy with access to meaningful, technically challenging work
  • Build systems at the intersection of software engineering and artificial intelligence
  • Collaborate with a global network of talented engineers and researchers
  • Potential for ongoing work and contract extension as new projects launch

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

Gerente IT

Gallagher

Sao Paulo

Apply on the employer's site

Role description

Introduction
Welcome to Gallagher - a global community of people who bring bold ideas, deep expertise, and a shared commitment to doing what’s right. We help clients navigate complexity with confidence by empowering businesses, communities, and individuals to thrive. At Gallagher, you’ll find more than a job; you’ll find a culture built on trust, driven by collaboration, and sustained by the belief that we’re better together. Whether you join us in a client-facing role or as part of our brokerage division, our benefits and HR consulting division, or our corporate team, you’ll have the opportunity to grow your career, make an impact, and be part of something bigger. Experience a workplace where you’re encouraged to be yourself, supported to succeed, and inspired to keep learning. That’s what it means to live The Gallagher Way.

Overview
As IT Manager for Brazil, you will lead technology operations for our São Paulo and Rio de Janeiro offices. Your work will keep people connected, systems secure, and the business running smoothly. This role matters because it supports teams every day and helps the company grow with reliable, scalable technology.

You will shape how IT supports the business (today and in the future) while leading people, partners, and platforms with care and purpose

How You'll Make An Impact

  • Ensure employees have reliable, secure, and well‑supported technology to do their best work.
  • Lead and support IT analysts, helping them grow and succeed as a team.
  • Drive system integrations and technology transitions aligned with global standards.
  • Translate complex technical needs into clear, practical guidance for business leaders.
  • Improve how IT works by creating clear processes, documentation, and support models.
  • Manage technology vendors and partnerships to deliver quality and value.
  • Strengthen data protection, cybersecurity, and compliance, including Brazil’s LGPD.
  • Help modernize infrastructure through migration strategies and new technical solutions.

About You

  • Experience leading IT teams and/or technical projects in a corporate environment.
  • Professional fluency in English.
  • Strong knowledge of Microsoft‑based environments, including Windows, Microsoft 365, Azure, Active Directory, and Exchange.
  • Experience with infrastructure, networks, telephony, virtual private networks (VPNs), and system integrations.
  • Confidence supporting and resolving complex third‑level (or higher) technical issues.
  • Experience working with Quiver or similar enterprise platforms.
  • The ability to explain technical concepts clearly to both technical and non‑technical audiences.
  • Experience managing vendors, equipment inventory, and IT services.
  • The ability to prioritize multiple requests while staying organized and calm.
  • A proactive, independent working style, with good judgment on when to escalate.

Nice to have, but not required:

  • Intermediate or advanced Spanish.
  • Project Management certification.
  • Experience in insurance, reinsurance, or brokerage environments.

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

34 more openings in this category and country

Site Reliability EngineerThales · Brazil

Apply on the employer's site
Site Reliability Engineer — Thales | mentors.coach