Skip to content
Site Reliability Engineer

The week's list

Every role like this one, in one letter

You are reading one posting. There are hundreds like it on the board, and new ones every week. Pick what you want, leave an email, and the list comes to you — no searching, no coming back here.

Counting what came out this past week…

The first letter arrives right away, then one a week. Unsubscribe in one click from any letter — the address goes nowhere else.

Site Reliability Engineer

Senior Manager, Production Engineering (EMEA)

CoreWeave

Warsawsenior

Apply on the employer's site

Role description

CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at www.coreweave.com.

We're proud to be a Living Wage accredited Employer.

What You'll Do
CoreWeave’s Production Engineering team builds, scales, and maintains the underlying infrastructure of the CoreWeave Cloud Platform. We operate at massive scale, partnering closely with downstream platform, hardware, and security teams to build secure, scalable, and highly efficient environments for demanding AI/ML and GPU-accelerated workloads. Our goal is to champion operational excellence, platform resilience, and automation-first infrastructure.

About The Role
As the Senior Manager of Production Engineering, you will lead, expand, and mentor a high-performing, geographically distributed SRE team operating in a 24x7 environment. You will own and execute the overarching SRE vision, strategy, and roadmap for our large-scale distributed cloud infrastructure, shifting the platform toward a proactive, self-healing architecture. This critical leadership role requires you to evolve our global on-call strategy, establish robust SLO/SLA frameworks, and champion structural improvements across incident management and postmortem organizational learning.

Who You Are

  • 5+ years of leadership or senior management experience at a cloud provider, hyperscaler, or high-growth technology company.
  • Proven experience hiring, developing, and managing geographically distributed, 24x7 engineering teams.
  • Extensive experience designing and implementing end-to-end incident management processes, including on-call rotations, escalation paths, blameless postmortems, and SLO/SLA frameworks.
  • Solid technical foundation in systems engineering, with a deep understanding of distributed systems, complex networking, and storage architecture.
  • Strong cross-functional collaboration and influencing skills, with the ability to align product, platform, hardware, and security teams on reliability goals.
  • A champion for automation-first practices, leveraging tools like Terraform, Kubernetes, and Infrastructure-as-Code to systematically eliminate manual toil.
  • A thoughtful leader who blends technical depth with strategic vision, preferring clarity over complexity, mentorship over management, and resilience over rigidity.
  • Bachelor’s degree in Computer Science, Engineering, or a related technical field.

Preferred

  • Experience building internal platform tooling and developer portals to enhance engineering velocity and operational visibility.
  • Familiarity with bare metal infrastructure environments (e.g., custom data centres, edge compute, or HPC clusters).
  • Experience working in AI infrastructure environments supporting large-scale training or inference pipelines.
  • Familiarity with GPU-accelerated workloads, including hardware resource isolation, scheduling, and performance tuning.
  • Background in compliance, reliability risk modelling, or operational maturity assessments (e.g., RTO/RPO limits, chaos engineering).
  • Working knowledge of DPUs, service mesh architectures, and multi-tenant security models.

Wondering if you're a good fit?
We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams even if you aren't a 100% skill or experience match. Here are a few qualities we've found compatible with our team. If some of this describes you, we'd love to talk.

  • You love to build clarity out of operational complexity and mentor high-performing SRE teams to scale distributed systems confidently.
  • You're curious about automation-first practices, infrastructure-as-code, and finding structural ways to eliminate manual toil through organizational learning.
  • You're an expert in managing incident lifecycles, driving blameless postmortem cultures, and building scalable, self-healing cloud architectures under pressure.

Why CoreWeave?
About
At CoreWeave, we work hard, have fun, and move fast! We're in an exciting stage of hyper-growth that you will not want to miss out on. We're not afraid of a little chaos, and we're constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values:

  • Be Curious at Your Core
  • Act Like an Owner
  • Empower Employees
  • Deliver Best-in-Class Client Experiences
  • Achieve More Together

We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and enables the development of innovative solutions to complex problems. As we get set for takeoff, the organisation's growth opportunities are constantly expanding. You will be surrounded by some of the best talent in the industry, who will want to learn from you, too. Come join us!

The typical cash compensation range for this position is
369,000–542,300 PLN
gross annually, with an additional performance-based bonus and equity component that can significantly increase total compensation up to ~1,500,000 PLN. The exact starting base salary will be determined by job-related knowledge, skills, experience, and the market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility).

To fulfill our obligation to protect client data, successful applicants offered employment with CoreWeave will be required to complete a basic criminal record check, conducted in compliance with GDPR. Employment offers are conditional upon receiving satisfactory check results.
What We Offer
In addition to a competitive salary, we offer a variety of benefits to support your needs, including:

  • Family-level Medical Insurance
  • Family-level Dental Insurance
  • Generous Pension Contribution
  • Life Assurance at 4x Salary
  • Critical Illness Cover
  • Employee Assistance Programme
  • Tuition Reimbursement
  • Work culture focused on innovative disruption

Benefits may vary by location.
Equal Opportunity
CoreWeave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace. All qualified applicants and candidates will receive consideration for employment without regard to race, color, religion, sex, disability, age, sexual orientation, gender identity, national origin, veteran status, or genetic information.
Recruitment Agencies
CoreWeave does not accept speculative CVs. Any unsolicited CVs received will be treated as the property of CoreWeave and your Terms & Conditions associated with the use of CVs will be considered null and void.
Any unsolicited CVs sent by your company to us – that is to say, in any situation where we have not directly engaged your company in writing to supply candidates for a specific vacancy – will be considered by us to be a “free gift”, leaving us liable for no fees whatsoever should we choose to contact the candidate directly and engage the candidate’s services, and will in no way establish any prior claim by your company to representation of that candidate should the candidate’s details also be submitted by any other party.
Export Control Compliance
This position requires access to export controlled information. To conform to U.S. Government export regulations applicable to that information, applicant must either be (A) a U.S. person, defined as a (i) U.S. citizen or national, (ii) U.S. lawful permanent resident (green card holder), (iii) refugee under 8 U.S.C.

  • 1157, or (iv) asylee under 8 U.S.C.
  • 1158, (B) eligible to access the export controlled information without a required export authorization, or (C) eligible and reasonably likely to obtain the required export authorization from the applicable U.S. government agency. CoreWeave may, for legitimate business reasons, decline to pursue any export licensing process.

Updated privacy notice - UK and EU Job Applications
When you apply to a job on this site, the personal data contained in your application will be collected by CoreWeave UK Ltd. (“Controller”), which is located at
Phosphor (6th Floor), 133 Park Street, London, SE1 9EA
and can be contacted by emailing
careers.eu@coreweave.com
. Controller’s data protection officer can be contacted at
privacy@coreweave.com
. Your personal data will be processed for the purposes of managing Controller’s recruitment related activities, which include setting up and conducting interviews and tests for applicants, evaluating and assessing the results thereto, and as is otherwise needed in the recruitment and hiring processes. Such processing is legally permissible under Art. 6(1)(f) of (i) Regulation (EU) 2016/679 (General Data Protection Regulation (“GDPR”) and (ii) the GDPR as it forms part of the laws of the UK (“UK GDPR”), as necessary for the purposes of the legitimate interests pursued by the Controller, which are the solicitation, evaluation, and selection of applicants for employment. Your personal data will be shared with Greenhouse Software, Inc., a cloud services provider located in the United States of America and engaged by Controller to help manage its recruitment and hiring process on Controller’s behalf. With respect to transfers originating from the UK or the European Economic Area ("EEA") to a country outside the UK or the EEA, we implement the appropriate transfer mechanism(s) and other appropriate solutions to address cross-border transfers as required by applicable law. You may request a copy of the suitable mechanisms we have in place by contacting us at
privacy@coreweave.com
Your personal data will be retained by Controller as long as Controller determines it is necessary to evaluate your application for employment. Where permitted by applicable law, we may also retain your personal data for a limited period after the recruitment process ends in order to consider you for future job opportunities, respond to legal claims, or comply with record-keeping obligations. Under the GDPR and the UK GDPR, you have the right to request access to your personal data, to request that your personal data be rectified or erased, and to request that processing of your personal data be restricted. You also have the right to data portability. In addition, you may lodge a complaint with the relevant supervisory authority: (i) A list of Europe’s data protection authorities can be found
here
; and (ii) for the UK, this is the
Information Commissioner's Office
.
For additional information, please see our
Privacy Policy
.

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

Technical Lead, Cluster Management System

Google

Warsawlead

Apply on the employer's site

Role description

Minimum qualifications:

  • Bachelor's degree in Computer Science, or a related technical field, or equivalent practical experience.
  • 5 years of experience in software architecture.

Preferred qualifications:

  • 2 year of experience in working with and influencing a group of people.
  • Coding experience in unmanaged language (Rust, C, C++).
  • Experience in concurrency, multithreading and synchronization.

About The Job
Google Cloud's software engineers develop the next-generation technologies that change how billions of users connect, explore, and interact with information and one another. We're looking for engineers who bring fresh ideas from all areas, including information retrieval, distributed computing, large-scale system design, networking and data storage, security, artificial intelligence, natural language processing, UI design and mobile; the list goes on and is growing every day. As a software engineer, you will work on a specific project critical to Google Cloud's needs with opportunities to switch teams and projects as you and our fast-paced business grow and evolve. You will anticipate our customer needs and be empowered to act like an owner, take action and innovate. We need our engineers to be versatile, display leadership qualities and be enthusiastic to take on new problems across the full-stack as we continue to push technology forward.

Google's Cluster Management System works on one of the most impactful software at Google, that affects all companies'workloads. We work on supporting the latest kinds workloads, latest hardware and latest security measures. Our mission is to improve Cluster Management System efficiency, make it easy to use, scalable and secure.

The Cluster Management System team runs as a group of small teams executing on various projects in all major areas of infrastructure. Each team is run independently and given a large amount of freedom, while still counting on the support of the rest of the team when necessary.

Google Cloud accelerates every organization’s ability to digitally transform its business and industry. We deliver enterprise-grade solutions that leverage Google’s technology, and tools that help developers build more sustainably. Customers in more than 200 countries and territories turn to Google Cloud as their trusted partner to enable growth and solve their most critical business problems.

Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

Poland: zł364000 - zł373000 (PLN) + 15% bonus target + equity + benefits

Responsibilities
Learn more about benefits at Google .

  • Understand the guest (virtualized) and host (hypervisor) environments and come up with secure solutions that connect both.
  • Work with peers to identify, design, create and optimize software features that run on top of Google's hardware stack.
  • Lead design and build processes for infrastructure to automatically measure and detect regressions in key performance metrics at scale.
  • Help address the trickiest security vulnerabilities.
  • Develop code in C++.

Google is proud to be an equal opportunity workplace and is an affirmative action employer. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements. See also Google's EEO Policy and EEO is the Law. If you have a disability or special need that requires accommodation, please let us know by completing our Accommodations for Applicants form .

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

STRUCTURAL DYNAMICS ENGINEER, FEM ANALYST - AVIATION INDUSTRY

Quest Global

Bielsko-Biała

Apply on the employer's site

Role description

Job Requirements

SHAPE THE FUTURE WITH US

We offer you a motivating environment with exciting projects and would appreciate to welcome you in our team as a: STRUCTURAL DYNAMICS ENGINEER, FEM ANALYST - AVIATION INDUSTRY

The selected resources will be inserted into an engineering support team in the Engine Dynamics discipline on development programs. They will be involved in design activities with the execution of dynamic structural analysis on components, engine modules and integrated engine systems, auxiliary systems, transmissions.

Work Experience

Key requirements:

  • Master's degree in Mechanical / Aeronautical / Aerospace Engineering;
  • At least one year of experience in FEM analysis and simulation (static, dynamic, thermal, etc.) of mechanical components, using tools such as Hypermesh, Nastran, ANSYS, etc. (professional knowledge of Nastran will be considered a plus);
  • Experience/ability to develop macros and scripts for post-processing activities;
  • Any background in turbomachinery and basic experience/ability in using CAD tools (e.g., Siemens NX, Solid Edge) will be considered a preferred requirement;
  • Good command of the English language.
  • The ideal profile is completed by strong analytical skills, problem-solving ability, organization and method, result-oriented mindset in a dynamic work context with tight deadlines, reliability, teamwork attitude, as well as strong communication and relationship management skills.

Benefits

  • Private healthcare (Medicover) – comprehensive private medical care.
  • Professional development – access to technical training and language courses to support your career growth.
  • Opportunity to work on innovative Aerospace & Aviation projects with leading global customers.
  • On-site collaboration with the client in a dynamic international engineering environment.
  • Salary range: our budget for this role is between PLN 11,000 - 12,000 gross per month.

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

Cloud Infrastructure Engineer

Ringier Axel Springer Polska

Krakowfulltime

Apply on the employer's site

Role description

Buduj platformę cloud dla ponad 40 milionów użytkowników

Masz dość projektów, które nigdy nie trafiają na produkcję? Chcesz pracować z AWS, Azure i GCP jednocześnie oraz rozwijać platformę wykorzystywaną przez setki inżynierów?

Dołącz do zespołu Cloud Infrastructure w Ringier Axel Springer Tech i twórz rozwiązania, które każdego dnia wspierają produkty używane przez ponad
40 milionów użytkowników na całym świecie
.

Dlaczego warto do nas dołączyć?

  • Realny wpływ: Tworzysz platformę cloud dla produktów wykorzystywanych codziennie przez ponad 40 milionów użytkowników. To infrastruktura działająca w skali enterprise, gdzie Twoje decyzje mają realny wpływ na biznes.
  • Multi-cloud w praktyce: Pracujemy jednocześnie z AWS, Azure i Google Cloud Platform. Nie ograniczasz się do jednej technologii — rozwijasz kompetencje w największych chmurach publicznych.
  • Autonomia i rozwój: Otrzymujesz dużą samodzielność, współpracujesz z doświadczonym zespołem i masz przestrzeń do proponowania własnych rozwiązań oraz eksperymentowania z nowymi technologiami.

Na Co Dzień Będziesz

  • projektować, rozwijać i utrzymywać infrastrukturę jako kod z wykorzystaniem Terraform, Terragrunt oraz CloudFormation,
  • rozwijać platformę cloud wspierającą pracę ponad 300 inżynierów,
  • budować rozwiązania zwiększające bezpieczeństwo oraz automatyzujące wykrywanie i usuwanie podatności,
  • tworzyć narzędzia do monitorowania kosztów infrastruktury i optymalizacji wykorzystania zasobów,
  • uczestniczyć w code review oraz dzielić się wiedzą z zespołem,
  • współpracować z zespołami produktowymi przy projektowaniu i wdrażaniu nowych rozwiązań,
  • rozwijać monitoring, polityki bezpieczeństwa oraz strategie Disaster Recovery,
  • tworzyć i aktualizować dokumentację techniczną,
  • aktywnie uczestniczyć w pracy zespołu Scrum.

Wymagania

  • AWS - znasz serwisy i wiesz, jak je używać w praktyce
  • Infrastructure as Code Terraform, CloudFormation
  • Python - do zastosowań budowy skryptów, biblioteki m.in. boto3
  • Git + GitHub - code review, CI/CD (GitHub actions)
  • "lifelong learner" mindset - technologie się zmieniają, Ty też powinieneś/aś się rozwijać.
  • Język angielski - min. B2

Nice to have:

  • Microsoft Azure - drugi kawałek multi-cloud puzzle
  • Google Cloud Platform - kompletny multi-cloud expert

Oferujemy

  • details .job--detail--content h2.header--h3.-medium.mb-5:nth-of-type(0n+3), .details .job--detail--content h2.header--h3.-medium.mb-5:nth-of-type(0n+3)+div { display: none !important }
  • No blame culture — szukamy rozwiązań, nie winnych
  • High autonomy — duża autonomia
  • Experienced developers to work with — doświadczeni programiści
  • Kultura DevOps
  • LeSS & Agile
  • Ubezpieczenie na życie dla ciebie i bliskich
  • Wsparcie psychologiczne i program well-being
  • Multisport package and sport challenges
  • Training budget & e-learning platforms
  • Legendarne imprezy
  • Prywatna opieka medyczna
  • AWS Certification Path — wsparcie w certyfikacji AWS
  • Nowoczesne, kolorowe i komfortowe biura
  • Transparentne widełki i ścieżki kariery
  • Lightning Talks, Code Jams & HackDays

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

Site Reliability Engineer, Mistral Cloud

Mistral

Warsawfulltime

Apply on the employer's site

Role description

About Mistral
Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector—co-creating customized AI systems that they can run on their terms.

We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited.

The Role
As a
Site Reliability Engineer (SRE)
on the
Cloud Platform
team, you will shape the reliability, scalability, and performance of our Cloud platform and customer-facing applications. You’ll work closely with software engineers and product teams to ensure our systems meet and exceed the expectations of both internal and external customers.

This role is critical in maintaining the stability and efficiency of our infrastructure, enabling seamless experiences for users and developers. Your expertise will directly impact the robustness of our AI platform, ensuring it operates at scale with minimal downtime.

What You Will Do

  • Design, build, and maintain scalable, highly available, and fault-tolerant infrastructures to support our Cloud platform.
  • Operate systems and troubleshoot issues in production environments, including interrupts, on-call responses, and infrastructure scaling.
  • Implement and improve monitoring, alerting, and incident response systems to minimize downtime and optimize performance.
  • Develop and maintain workflows and tools for CI/CD, containerization, orchestration, monitoring, and logging.
  • Participate in on-call rotations to respond to incidents and perform root cause analysis to prevent recurrence.
  • Drive continuous improvement in infrastructure automation, deployment, and orchestration.
  • Collaborate with software engineers to enable safe and reproducible model-training experiments.
  • Build and enhance a cloud platform that abstracts infrastructure complexities for science and engineering teams.
  • Design and develop new workflows, tooling, and automation to improve system reliability, availability, and performance.
  • Ensure infrastructure adheres to security best practices and compliance requirements in collaboration with the security team.
  • Document processes and procedures to ensure consistency and knowledge sharing across the team.

What We're Looking For

  • A Master’s degree in Computer Science, Engineering, or a related field.
  • 5+ years of experience in a DevOps or SRE role, with strong expertise in bare metal infrastructure and distributed systems.
  • Hands-on experience with site reliability issues, including root cause analysis, in-production troubleshooting, and on-call rotations.
  • Proficiency in working with reliability KPIs, such as observability, alerting, and SLAs.
  • Experience with CI/CD, containerization, and orchestration tools like Docker and Kubernetes.
  • Knowledge of monitoring, logging, alerting, and observability tools such as Prometheus, Grafana, ELK Stack, or Datadog.
  • Familiarity with infrastructure-as-code tools like Terraform or CloudFormation.
  • Proficiency in scripting languages (Python, Go, Bash) and a strong understanding of software development best practices.
  • Solid grasp of networking, security, and system administration concepts.
  • Excellent problem-solving and communication skills, with the ability to work effectively in a collaborative environment.
  • Experience in an AI/ML environment, high-performance computing (HPC) systems, or modern AI-oriented solutions (e.g., Fluidstack, Coreweave, Vast) is a plus.

What we offer
We offer a comprehensive benefits package designed to support your well-being, growth, and work-life balance. Benefits vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks.

For the most up-to-date details on benefits available in your location, please refer to our Benefits page.

Privacy Policy
Your privacy matters to us. You can learn more about how we handle your personal data in our Applicant Privacy Policy.

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

85 more openings in this category and country

Senior Manager, Production Engineering (EMEA)CoreWeave · Poland

Apply on the employer's site