Skip to content
Site Reliability Engineer

Site Reliability Engineer

Coordenador(a) de Suporte

Starbem

Sao Paulo

Откликнуться на сайте работодателя

Описание вакансии

Descrição da vaga

Procuramos alguém que já coordenou operações de suporte estruturadas em níveis (N1, N2 e N3) e que saiba fazer as três coisas ao mesmo tempo: liderar gente, operar com método (ITIL/ITSM) e usar dado para decidir.Essa pessoa será dona da operação de suporte ponta a ponta — do primeiro contato ao encerramento, passando por escalonamento, gestão de incidentes, base de conhecimento e melhoria contínua. Também será a interface formal entre Suporte, Engenharia, Produto e Segurança.

Responsabilidades e atribuições

O que buscamos, em uma frase: alguém curioso o suficiente para investigar a causa raiz em vez de fechar o ticket, e disciplinado o suficiente para transformar essa investigação em processo.ResponsabilidadesOperação e método (ITIL/ITSM)

  • Estruturar e manter a operação em três níveis: definir claramente o que resolve em N1, o que sobe para N2 e o que chega em N3, com critérios objetivos de escalonamento.
  • Aplicar as práticas de ITIL v4 no dia a dia: gestão de incidentes, requisições de serviço, gestão de problemas, gestão de mudanças e catálogo de serviços.
  • Definir, publicar e defender SLAs e SLOs por categoria e severidade — incluindo o que acontece quando o SLA é estourado.
  • Coordenar a resposta a incidentes críticos (major incidents): comando da sala, comunicação com stakeholders, atualização de status e encerramento formal.
  • Conduzir post-mortems sem culpados, com plano de ação rastreável, e cobrar o fechamento das ações com Engenharia e Produto.
  • Manter runbooks, playbooks de triagem e base de conhecimento vivos — documentação desatualizada é dívida operacional.
  • Gerir a fila e o backlog: priorização, redistribuição, prevenção de ticket envelhecido e controle de reincidência.

Liderança de pessoas

  • Liderar, desenvolver e avaliar os times de N1, N2 e N3 — 1:1s regulares, feedback direto, plano de desenvolvimento individual e trilha de carreira entre níveis.
  • Dimensionar o time e a escala de atendimento com base em volume e sazonalidade; propor headcount com justificativa quantitativa.
  • Estruturar onboarding, ramp-up e certificação de novos analistas.
  • Rodar garantia de qualidade (QA) de atendimentos com rubrica clara, e usar o resultado para treinar, não para punir.
  • Organizar escala de plantão/sobreaviso quando aplicável, com regras claras de acionamento e de descanso.

Dados e melhoria contínua

  • Ser dono das métricas da operação e reportá-las com frequência fixa para a liderança (ver quadro de KPIs abaixo).
  • Identificar os principais geradores de contato e atacar a causa: se o mesmo problema gera 200 tickets por mês, a solução está no produto, não no roteiro de atendimento.
  • Levar para Produto e Engenharia uma pauta priorizada por impacto real, com evidência — não com percepção.
  • Avaliar, implantar e administrar as ferramentas de service desk (ex.: Jira Service Management, Zendesk, Freshdesk) e sua integração com os demais sistemas.

Interfaces e conformidade

  • Ser o ponto de contato formal do Suporte com Engenharia, Produto, Segurança da Informação, Jurídico e a operação de saúde.
  • Garantir que a operação trate dados pessoais e dados de saúde conforme a LGPD: minimização, controle de acesso, registro de tratamento e rota clara para incidentes de dados.
  • Assegurar que o time nunca dê orientação clínica: temas de saúde são encaminhados a profissionais habilitados e situações de urgência seguem o protocolo de emergência definido pela operação de saúde.

Requisitos e qualificações

  • Experiência comprovada coordenando ou gerenciando equipes de suporte estruturadas em N1, N2 e N3 — liderando pessoas diretamente, não apenas atuando como analista sênior.
  • ITIL / ITSM na prática — gestão de incidentes, problemas, mudanças e requisições; certificação ITIL v4 Foundation é um diferencial forte, mas o que vale é a prática.
  • Domínio de SLA e severidade — desenhar, negociar, medir e defender acordos de nível de serviço.
  • Gestão de incidentes críticos — já esteve à frente de uma indisponibilidade relevante, com comunicação a stakeholders sob pressão.
  • Operação orientada a dados — constrói e lê os próprios indicadores, e toma decisão com base neles.
  • Ferramentas de service desk — administração (não só uso) de pelo menos uma plataforma de ITSM, incluindo automações, filas, formulários e relatórios.
  • Comunicação escrita clara em português — comunicado de incidente, post-mortem e documentação de processo são entregáveis dessa vaga.
  • Consciência de LGPD e confidencialidade — a operação lida com dados de saúde; sigilo não é opcional.

Requisitos desejáveis

Camada de segurança

  • Vivência com gestão de acessos e identidade (IAM, SSO, MFA, princípio do menor privilégio, revisão periódica de acessos).
  • Participação em resposta a incidentes de segurança e no fluxo de comunicação de incidente de dados.
  • Familiaridade com controles de segurança em endpoints e com processos de conformidade (ex.: ISO 27001, SOC 2) — mesmo que como área participante, não como dono.
  • Noção de segurança aplicada ao próprio suporte: verificação de identidade do solicitante, prevenção de engenharia social e cuidado com dados sensíveis em ticket.

Inteligência artificial aplicada a suporte corporativo

  • Experiência implantando IA em operação de suporte: assistentes de atendimento, sugestão de resposta, triagem e classificação automática, sumarização de ticket, busca semântica na base de conhecimento.
  • Entendimento prático dos limites da tecnologia: quando a IA resolve, quando ela deve entregar para um humano, e como medir a qualidade da resposta gerada.
  • Cuidado com dados sensíveis em fluxos com IA: o que pode e o que não pode ser enviado a um modelo, e por quê.
  • Capacidade de avaliar ganho real (deflexão, tempo de resolução, custo por ticket) em vez de adotar ferramenta por moda.

Outros diferenciais

  • Familiaridade com ferramentas de observabilidade (ex.: Datadog, Grafana, New Relic) para triagem técnica com base em logs, traces e alertas.
  • SQL básico e leitura de painéis de dados para investigar sem depender de terceiros.
  • Experiência em healthtech, saúde suplementar ou em ambientes regulados.
  • Inglês técnico para leitura de documentação e interação com fornecedores.
  • Noção de arquitetura de sistemas modernos (APIs, integrações, aplicativos móveis) o suficiente para conversar de igual para igual com Engenharia.

Enviar candidatura

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

Скоро на этой странице

Резюме под эту вакансию — и билет в розыгрыш

Мы разбираем объявление до настоящих требований и переписываем ваше резюме под него — вопросами, а не выдумкой: ни одна строка не появится без вашего подтверждения. Войдите, чтобы получить это первым, — и попасть в розыгрыш.

  • Резюме под конкретную вакансию, а не «универсальное»
  • Ответы хранятся: правится любой, а не весь разговор заново
  • Всё в аккаунте — открывается с любого устройства

Разыгрываем

Скидка на сопровождение

Победителей выбираем случайно среди заявок с подтверждённой почтой. Дата розыгрыша и полные правила — на странице розыгрыша.

Правила розыгрыша

Site Reliability Engineer

Desenvolvedor Fullstack (Rancher) SR

AVANTTi

Sao Paulo

Откликнуться на сайте работодателя

Описание вакансии

  • Contratante: Avantti Consultoria TI
  • Atuação: 100% Remota
  • Contrato: CLT + Benefícios

Na Avantti Consultoria TI buscamos profissionais sêniores que queiram ir além do desenvolvimento convencional. Você fará parte de uma squad técnica com foco em
engenharia de software
de ponta, atuando em ambientes complexos e desafiadores, com metodologias ágeis e práticas modernas de mercado.

Nosso objetivo é garantir resiliência tecnológica, evolução contínua dos sistemas e prevenção de falhas, sempre alinhados à cultura SRE e às melhores práticas de DevSecOps.

Requisitos
🔹
Responsabilidades

  • Atuar como Desenvolvedor Full Stack em projetos técnicos.
  • Realizar análise crítica de código, propor arquitetura sistêmica e selecionar soluções modernas em ambientes TST, HML e PRD.
  • Desenvolver, refatorar e evoluir sistemas em Java (Spring Boot, Spring Batch, OpenLiberty) e Angular.
  • Garantir resolução de bugs na causa raiz.
  • Assegurar saúde e resiliência dos ambientes de produção via observabilidade, tuning de JVM e troubleshooting em clusters Kubernetes geridos pelo Rancher.
  • Atuar em prevenção/predição de falhas (SRE), gestão de capacidade (Requests/Limits), saúde de containers e expurgo otimizado de bases de dados.
  • Conduzir ações de obsolescência sistêmica, mitigação/correção de vulnerabilidades (AppSec) e otimização de imagens Docker em pipelines CI/CD.
  • Introduzir ferramentas de Inteligência Artificial (IA) para acelerar desenvolvimento, refatoração e automação de testes.

🔹
Requisitos

  • Experiência sólida em análise de código, refatoração e arquitetura sistêmica.
  • Domínio em Java (Back-End) e Angular/TypeScript (Front-End), com suporte a legado (JavaScript/jQuery/HTML).
  • Vivência com Spring Boot, Spring Batch, OpenLiberty/Jakarta EE.
  • Experiência prática com Docker, Kubernetes e Rancher (troubleshooting, monitoramento e governança).
  • Domínio em microsserviços, APIs REST e padrões de resiliência (Circuit Breakers, Retries, Graceful Shutdown, Rate Limiting).
  • Conhecimento avançado em Oracle/SQL (queries, procedures, performance, expurgo).
  • Experiência com ferramentas de CI/CD e DevSecOps: Git, GitLab, Jenkins, SonarQube, Trivy, Snyk.
  • Conhecimento prático em cultura SRE, metodologias ágeis (Scrum/Kanban) e ITIL.

🔹
Desejáveis

  • Certificações: Kubernetes CKAD/CKA, Java, Spring, Rancher, DevSecOps.
  • Conhecimento em DataStage (modelagem e engenharia de dados).
  • Experiência com Cloud (AWS, Azure, GCP).
  • Vivência com IA aplicada à engenharia de software (Copilot, geração/análise de código).

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

Site Reliability Engineer

Support and Monitoring Analyst

IRiS Grupo Tecnológico

Sao Paulo

Откликнуться на сайте работодателя

Описание вакансии

Are you passionate about cloud infrastructure and looking to be part of a dynamic team? At uCloud, we are looking for a Cloud Support & SysOps Analyst to join our global operations and monitoring team.

As part of your role, you will be responsible for:

  • Critical Monitoring & Operations: Ensure the availability and performance of cloud infrastructure by strictly following and executing documented processes (Runbooks/Playbooks).
  • Immediate Incident Response: Act with agility and dynamism in response to critical monitoring alerts and urgent infrastructure technical support requests.
  • GCP Environment Support: Provide operational support for core services such as Compute Engine, Google Kubernetes Engine (GKE), Identity and Access Management (IAM), and storage administration in Google Cloud Storage.
  • Continuous Improvement & Automation: Adopt a proactive approach to suggest improvements to the technical environment, optimize observability dashboards, and propose automations for repetitive tasks.
  • User & Workspace Support: Administer the Google Admin console and associated corporate tools to resolve incidents with a strong focus on User Experience (UX).
  • Operational Management (ITSM): Register, classify, and document the lifecycle of incidents and requests using Jira Service Manager.

Requisitos mínimos

  • Cloud Computing: Strong hands-on experience with Google Cloud Platform (GCP).
  • Cost Estimation: Ability to create budgets and projections using the Google Cloud Pricing Calculator.
  • Infrastructure & Log Analysis: Ability to perform cloud infrastructure and log analysis using native Google tools.
  • Monitoring & Observability: Consolidated experience with Zabbix and GCP tools: Google Cloud Monitoring (including implementing alert policies), configuring Uptime Checks, and advanced analysis via Logs Explorer.
  • GCP Infrastructure & Storage: Experience with VM instances (Compute Engine), GKE (Google Kubernetes Engine), permission management (IAM), database infrastructure support (Cloud SQL — without internal administration focus), and Bucket management in Google Cloud Storage (understanding classes such as Standard, Nearline, Coldline, and Archive).
  • Google Workspace Ecosystem: Operational mastery of corporate tools (Gmail, Chat, Calendar, Meet, Sheets, Docs, etc.) and experience with Google Admin to provide technical support to users, alongside daily work usage.
  • ITSM Tools: Day-to-day experience using Jira Service Manager (Atlassian Suite) or a similar tool.
  • Customer Service: Previous experience in technical support focused on User Experience (UX) and agile problem-solving.

Desirable & Differential Requirements

  • Infrastructure as Code & CLI: Desirable knowledge of Terraform and proficiency in using Cloud Shell.
  • FinOps & Cost Management: Experience in billing analysis and billing permission management, as well as consumption control through Budget Alerts configurations.
  • Cloud & Workspace Certifications:
  • Active Cloud Digital Leader or Google Cloud Certified - Associate Cloud Engineer certifications.
  • Certifications oriented toward Google Workspace administration (Major differential).
  • Data Analysis & Reporting: Ability to design operational reports and dashboards using Looker Studio (formerly Data Studio).
  • Security & Governance: Practical knowledge of secret management with Secret Manager and vulnerability monitoring through Security Command Center.
  • Serverless Architecture & Messaging: Knowledge or experience with deployments in Cloud Run and queue/event management with Pub/Sub.
  • Operating Systems: Technical knowledge in Linux environments (equivalent to LPIC-1 and LPIC-2 certification requirements).
  • Automation & Development: Basic knowledge of programming logic (in any scripting language) oriented toward automating daily processes.
  • Documentation & AI: Experience drafting internal technical documentation and using Artificial Intelligence tools to optimize routines.

Languages

  • Portuguese: Native or advanced fluency (mandatory for reading, writing, and conversation in daily internal alignments).
  • English: Intermediate reading comprehension level (focus on reading and interpreting technical IT documentation).
  • Spanish: Fluid or advanced level (completely optional/desirable, considered solely as a differential for regional communication).
  • Schedule & Shift: 12-hour alternating night shift (Standard from 18:00 to 06:00).
  • Important Note: Availability for coverage/swapping with the day shift (06:00 to 19:00) up to 2 times a week is mandatory, preferably on Thursdays and Fridays.

Benefits

  • PJ Contract
  • Health Insurance

#Ucloud

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

Site Reliability Engineer

Back End Developer (AI Infrastructure)

Alignerr

Sao Paulo

Откликнуться на сайте работодателя

Описание вакансии

About The Role
What if the code you write could directly power the AI systems shaping how millions of people interact with technology? We're looking for Back End Developers in São Paulo to design, build, and optimize the server-side systems and APIs that drive cutting-edge AI products — from data pipelines and model-serving infrastructure to scalable microservices that handle real-world traffic at scale.

This is a fully remote, flexible contract role. Whether you're a seasoned engineer looking for meaningful project work or a strong developer ready to apply your skills to AI-adjacent infrastructure, this is your chance to build things that matter.

  • Organization: Alignerr
  • Type: Hourly Contract
  • Location: Remote
  • Commitment: 10–40 hours/week

What You'll Do

  • Design, develop, and maintain robust back end services, APIs, and data pipelines
  • Build scalable, high-performance server-side architecture to support AI training and inference workloads
  • Write clean, well-documented, and testable code following engineering best practices
  • Integrate with databases, message queues, cloud services, and third-party APIs
  • Optimize system performance, reliability, and scalability as usage grows
  • Collaborate asynchronously with distributed teams of engineers, researchers, and product leads
  • Troubleshoot production issues, conduct code reviews, and contribute to architectural decisions
  • Work independently on task-based and project-based assignments on your own schedule

Who You Are

  • Proficient in one or more back end languages — Python, Go, Java, Node.js, Rust, or similar
  • Solid understanding of RESTful API design, microservices architecture, and distributed systems
  • Experienced with relational and/or NoSQL databases (PostgreSQL, MongoDB, Redis, etc.)
  • Comfortable working with cloud platforms such as AWS, GCP, or Azure
  • Familiar with version control (Git), CI/CD pipelines, and containerization (Docker, Kubernetes)
  • Strong problem-solving skills and the ability to debug complex systems independently
  • Self-motivated, reliable, and effective when working asynchronously without close supervision
  • Clear communicator who can document decisions and collaborate across time zones

Nice to Have

  • Experience building infrastructure for machine learning or AI systems
  • Familiarity with data processing frameworks (Apache Kafka, Spark, Airflow, or similar)
  • Background in performance optimization, caching strategies, or high-throughput systems
  • Knowledge of authentication, authorization, and security best practices
  • Experience with serverless architectures or event-driven design patterns
  • Contributions to open-source projects or a strong portfolio of personal/professional work

Why Join Us

  • Work on cutting-edge AI infrastructure projects alongside leading research labs
  • Fully remote and flexible — set your own hours and work from anywhere
  • Freelance autonomy with access to meaningful, technically challenging work
  • Build systems at the intersection of software engineering and artificial intelligence
  • Collaborate with a global network of talented engineers and researchers
  • Potential for ongoing work and contract extension as new projects launch

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

Site Reliability Engineer

Gerente IT

Gallagher

Sao Paulo

Откликнуться на сайте работодателя

Описание вакансии

Introduction
Welcome to Gallagher - a global community of people who bring bold ideas, deep expertise, and a shared commitment to doing what’s right. We help clients navigate complexity with confidence by empowering businesses, communities, and individuals to thrive. At Gallagher, you’ll find more than a job; you’ll find a culture built on trust, driven by collaboration, and sustained by the belief that we’re better together. Whether you join us in a client-facing role or as part of our brokerage division, our benefits and HR consulting division, or our corporate team, you’ll have the opportunity to grow your career, make an impact, and be part of something bigger. Experience a workplace where you’re encouraged to be yourself, supported to succeed, and inspired to keep learning. That’s what it means to live The Gallagher Way.

Overview
As IT Manager for Brazil, you will lead technology operations for our São Paulo and Rio de Janeiro offices. Your work will keep people connected, systems secure, and the business running smoothly. This role matters because it supports teams every day and helps the company grow with reliable, scalable technology.

You will shape how IT supports the business (today and in the future) while leading people, partners, and platforms with care and purpose

How You'll Make An Impact

  • Ensure employees have reliable, secure, and well‑supported technology to do their best work.
  • Lead and support IT analysts, helping them grow and succeed as a team.
  • Drive system integrations and technology transitions aligned with global standards.
  • Translate complex technical needs into clear, practical guidance for business leaders.
  • Improve how IT works by creating clear processes, documentation, and support models.
  • Manage technology vendors and partnerships to deliver quality and value.
  • Strengthen data protection, cybersecurity, and compliance, including Brazil’s LGPD.
  • Help modernize infrastructure through migration strategies and new technical solutions.

About You

  • Experience leading IT teams and/or technical projects in a corporate environment.
  • Professional fluency in English.
  • Strong knowledge of Microsoft‑based environments, including Windows, Microsoft 365, Azure, Active Directory, and Exchange.
  • Experience with infrastructure, networks, telephony, virtual private networks (VPNs), and system integrations.
  • Confidence supporting and resolving complex third‑level (or higher) technical issues.
  • Experience working with Quiver or similar enterprise platforms.
  • The ability to explain technical concepts clearly to both technical and non‑technical audiences.
  • Experience managing vendors, equipment inventory, and IT services.
  • The ability to prioritize multiple requests while staying organized and calm.
  • A proactive, independent working style, with good judgment on when to escalate.

Nice to have, but not required:

  • Intermediate or advanced Spanish.
  • Project Management certification.
  • Experience in insurance, reinsurance, or brokerage environments.

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

Ещё 35 вакансий по этой категории в этой стране

Coordenador(a) de SuporteStarbem · Brazil

Откликнуться на сайте работодателя