Skip to content
Site Reliability Engineer

Site Reliability Engineer

Software Engineering Manager, Site Reliability Engineering, Gemini Enterprise Agent Platform, Google Cloud

Google

Warsawfulltimelead

Apply on the employer's site

Role description

Minimum qualifications:

  • Bachelor’s degree in Computer Science, a related field, or equivalent practical experience.
  • 8 years of experience with software development in one or more programming languages.
  • 3 years of experience managing people or teams.
  • 3 years of experience leading projects.
  • 3 years of experience designing, analyzing, and troubleshooting distributed systems.

Preferred qualifications:

  • Master's degree in Computer Science or Engineering.
  • Experience working in an agile software development environment.
  • Familiarity with Site Reliability Engineering (SRE) or Production Engineering practices, such as managing system reliability, uptime, and performance, as well as on-call and incident-response.
  • Familiarity with high-scale container orchestration fleet operations (e.g., GKE/Kubernetes) or specialized AI/ML compute infrastructure management involving GPUs, TPUs, or model-serving performance optimization.
  • Familiarity with security engineering, platform hardening, or data privacy protocols in large-scale environments.

About The Job
Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that Google's services—both our internally critical and our externally-visible systems—have reliability, uptime appropriate to users' needs and a fast rate of improvement. Additionally SRE’s will keep an ever-watchful eye on our systems capacity and performance.

Much of our software development focuses on optimizing existing systems, building infrastructure and eliminating work through automation. On the SRE team, you’ll have the opportunity to manage the complex challenges of scale which are unique to Google, while using your expertise in coding, algorithms, complexity analysis and large-scale system design.

SRE's culture of intellectual curiosity, problem solving and openness is key to its success. Our organization brings together people with a wide variety of backgrounds, experiences and perspectives. We encourage them to collaborate, think big and take risks in a blame-free environment. We promote self-direction to work on meaningful projects, while we also strive to create an environment that provides the support and mentorship needed to learn and grow.

To learn more:
check out our books on Site Reliability Engineering or read a career profile about why a Software Engineer chose to join SRE.

Gemini Enterprise supports platforms for third-party models, including popular open-source models from Model Garden and enterprise-grade models from key partners like Anthropic.

In this role, you will be at the forefront of ensuring the rock-solid reliability and performance of the infrastructure underpinning these third-party model deployments, with a focus on our GKE-based serving stack.

Behind everything our users see online is the architecture built by the Technical Infrastructure team to keep it running. From developing and maintaining our data centers to building the next generation of Google platforms, we make Google's product portfolio possible. We're proud to be our engineers' engineers and love voiding warranties by taking things apart so we can rebuild them. We keep our networks up and running, ensuring our users have the best and fastest experience possible.Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

Poland: zł480000 - zł492000 (PLN) + 20% bonus target + equity + benefits

Responsibilities
Learn more about benefits at Google .

  • Lead a team of site reliability engineers (SREs), guide their professional development, and ensure their success.
  • Own end-to-end availability and performance of Gemini Enterprise Agent Platform, and build automation to prevent problem recurrence. Automate response to all non-exceptional service conditions.
  • Lead by example, mentor the team, and establish credibility through quality technical execution, including code reviews.
  • Extend site reliability engineering (SRE) best practices across the Cloud AI development organization to ensure operational reliability at scale.
  • Take part in and manage on-call rotations across continents, using a follow-the-sun model.

Google is proud to be an equal opportunity workplace and is an affirmative action employer. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements. See also Google's EEO Policy and EEO is the Law. If you have a disability or special need that requires accommodation, please let us know by completing our Accommodations for Applicants form .

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Coming to this page

A resume for this role — and a ticket to the draw

We take the posting apart down to the real requirements and rewrite your resume against it — by asking, not inventing: no line appears without your confirmation. Sign in to get it first, and to enter the draw.

  • A resume for this exact role, not a universal one
  • Answers are kept: edit any one, not the whole conversation
  • All in your account — open it from any device

On the wheel

A discount on mentoring

Winners are drawn at random among entries with a confirmed email. The date and the full rules are on the draw page.

Draw rules

Site Reliability Engineer

Senior Cloud Engineer (Blockchain) - Digital Assets

EPAM Systems

Krakowcontractsenior

Apply on the employer's site

Role description

We are looking for a
Senior Cloud Engineer
to join our Digital Assets Crew, working on pioneering blockchain-based projects including tokenization of financial assets. Our team operates within two Scrum teams and builds cloud-hosted applications that leverage cutting-edge technologies to deliver value and innovation to the finance industry

Responsibilities

  • Develop cloud solutions that scale and ensure high availability
  • Nurture a DevOps culture across the team
  • Contribute to DLT-related projects, particularly blockchain
  • Design and maintain infrastructure using Terraform and Helm
  • Implement GitOps and CI/CD pipelines
  • Monitor distributed systems using Grafana and Prometheus
  • Manage secret management solutions across cloud environments
  • Collaborate with colleagues in English on cross-functional initiatives

Requirements

  • 5+ years of experience in DevOps/SRE roles
  • Expertise in Azure Cloud and Kubernetes
  • Proficiency in Terraform, Helm and GitOps/CI-CD
  • Experience with Grafana, Prometheus and distributed systems architecture at enterprise scale
  • Knowledge of secret management solutions
  • Good communication skills, comfortable interacting with colleagues in English (B2+)

Nice to have

  • Experience with Istio Service Mesh and zero-trust architecture
  • Background in banking or regulated environments
  • Familiarity with blockchain technology

We offer

  • We gather like-minded people:
  • Top tech minds driving innovation in AI, cloud and digital platform modernization
  • Supportive team and agile, startup-like culture
  • Hybrid by design mode and opportunity to work remotely within Poland
  • Chance to work abroad for up to 60 days annually
  • Business-driven relocation opportunities
  • We provide growth opportunities:
  • Career development programs
  • Thought leadership, mentoring, soft skills and well-being programs
  • Certification (Anthropic, Gemini, GCP, Azure, AWS)
  • English classes
  • We cover it all:
  • Stable pay
  • Participation in the Employee Stock Purchase Plan with a 15% discount
  • Benefits package (health insurance, multisport, shopping vouchers)
  • Referral bonuses up to $2,000
  • Offices featuring entertainment and relaxation zones, table tennis and football, free snacks, coffee and more
  • Corporate, social and well-being events
  • Please, note:
  • Benefits listed above are available to employees only
  • We are open for working with Contractors. Terms of B2B cooperation agreements are agreed individually
  • We will reach out to selected candidates exclusively

EPAM is global leader in AI transformation engineering and integrated consulting, serving Forbes Global 2000 companies and ambitious startups. With over thirty years of expertise in custom software, product and platform engineering, we empower our clients to become AI-Native enterprises, driving measurable value from innovation and digital investments.

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

ITIL Service Management Expert (EU Institution) - hybrid in Poland

NRB

Warsaw

Apply on the employer's site

Role description

Job Description
Who are we?
KEYES
is a dynamic global organization that takes pride in being the trusted partner of
EU Institutions.
With strong commitment to excellence and a
30-years track record
of delivering high-quality solutions, we are dedicated to supporting the growth and success of our clients. Our Mission is to help our clients keep up with the challenges of
digital transformation
by providing the right talent at the right time for the right job. To this end, we are constantly looking for talented professionals who are interested in working on
challenging international projects
and able to deliver high-quality results within multicultural environments. Our services include (but are not limited to)
modernization of solutions, digital workspaces, cloud technologies and IT security
. Our Headquarters are in Brussels and we have active accounts and offices across Europe (i.e. Luxembourg, Amsterdam, Athens, Stockholm, Geneva).

Is this YOU?
For our customer based in
Warsaw, Poland -
an European Institution, we are looking for
ITIL Service Manager Expert based in Poland
to join a long-term mission in the area of
cybersecurity, public sector
and
law enforcement
. You will play a key role in managing business architecture, guiding solution designs, and fostering collaboration among business architects and project teams to ensure coherence with the customer’s long-term vision.

Please note that the role requires 80-100% hybrid presence at the client’s headquarters in Warsaw, meaning that being based 2-3 hours by ground transportation from Warsaw will be required.
More specifically, you will be responsible for…

  • Supporting execution of the existing IT service management processes
  • Analysing existing IT service management processes
  • Identifying risks and weak points and recommending improvements
  • Advising on the best IT service management practices
  • Preparing, documenting and implementing improvement plans
  • Supporting implementation and customisation of ITSM solution
  • Participation in projects related with implementation and customization of ITSM solution
  • Preparing required policies and Standard Operating Procedures (SOPs) supporting IT service management processes
  • Creating a positive customer experience
  • Other specific duties as assigned by the ICT SCSMT team leader.

Job Requirements
Are you the perfect match?

  • University degree (BSc/MSc)
  • At least 3 years (full time) experience at the similar position
  • At least 2 years (full time) experience in managing ICT service delivery and/or ICT operations
  • At least ITIL Intermediate certificate (ITIL Expert certificate preferred)
  • Excellent practical knowledge of IT Service Management
  • Practical experience with ITSM solutions
  • Well understanding of complex information systems and their interoperability
  • Excellent understanding and practical knowledge of IT technologies and information systems technical components
  • Familiarity with project management approaches, tools and phases of the project lifecycle Skills:
  • Excellent communication skills (in written and verbal communication)
  • Very strong sense of responsibility
  • Accuracy and attention to details
  • Very good organizational skills
  • Very good reporting skills
  • Supportive and helpful personality with co-operative and service oriented attitude
  • Forward looking with a holistic approach
  • High level of motivation and initiative.

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

IT/OT Systems Engineer

ALGOTEQUE Innovation Hub

Warsawsenior

Apply on the employer's site

Role description

ALGOTEQUE is an IT consultancy firm that helps startups, mid-sized and large corporations to create and deliver innovative technologies.

Our team has a successful track record in designing, developing, implementing, and integrating software solutions (AI, ML, BI, Web, Automation) for Telecom, Energy, Bank, Insurance, Pharma, Automotive, Industry, e-commerce. We deliver our services both in fixed-price and time-and-materials models, helping our customers achieve their business and IT strategies.

Job Description
Poszukujemy Inżyniera Systemów IT/OT, który wesprze rozwój i utrzymanie infrastruktury technologicznej zakładu produkcyjnego. Konsultant będzie kluczowym ogniwem łączącym świat IT (infrastruktura, sieci, systemy) z OT (automatyka, linie produkcyjne, systemy sterowania), wspierając inicjatywy z obszaru Industry 4.0.

Zakres Projektu

  • Utrzymanie ciągłości działania systemów produkcyjnych (SCADA, HMI, MES)
  • Integracja systemów OT z infrastrukturą IT (serwery, sieci, środowiska chmurowe)
  • Wsparcie cyfryzacji procesów produkcyjnych i wdrożeń Industry 4.0
  • Diagnostyka i rozwiązywanie problemów na styku IT/OT (incydenty, wydajność)
  • Zarządzanie infrastrukturą sieciową w środowisku produkcyjnym (VLAN, segmentacja)
  • Współpraca z działami utrzymania ruchu, automatykami i zespołami IT
  • Wdrażanie i utrzymanie standardów bezpieczeństwa w środowisku OT
  • Tworzenie dokumentacji technicznej oraz procedur operacyjnych

Profile / Requirements

  • 3–5 lat doświadczenia w środowiskach IT/OT lub infrastrukturalnych
  • Praktyczna znajomość systemów przemysłowych (SCADA, MES, PLC – mile widziane)
  • Doświadczenie w pracy w środowisku produkcyjnym (fabryka / zakład przemysłowy)
  • Znajomość sieci przemysłowych i IT (TCP/IP, VLAN, routing, firewall)
  • Systemy operacyjne: Windows Server i/lub Linux
  • Umiejętność pracy w środowisku krytycznym (wysoka dostępność, uptime)
  • Język angielski min. B2

Mile Widziane

  • Znajomość standardów bezpieczeństwa OT (ISA/IEC 62443)
  • Doświadczenie z systemami MES / integracją danych produkcyjnych
  • Chmura (Azure / AWS) w kontekście zbierania i analizy danych z produkcji
  • Podstawy automatyki (PLC – Siemens, Rockwell)
  • Skrypty / automatyzacja (PowerShell, Python)

AO4344

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

Administratorka/Administrator aplikacji

PKO Bank Polski

Warsaw

Apply on the employer's site

Role description

Na Co Dzień w Naszym Zespole

  • administrujemy aplikacjami z obszarów zarządzania tożsamością, HR oraz gospodarki własnej Banku,
  • utrzymujemy dostępność aplikacji administrowanych przez nasz zespół poprzez ciągłe monitorowanie zasobów jak i procesów, wgrywanie zmian oraz poprawek aplikacyjnych,
  • realizujemy usługi informatyczne dla jednostek biznesowych banku w zakresie eksploatowanych aplikacji,
  • rozwiązujemy incydenty/problemy zgłaszane przez jednostki banku z zakresu działania administrowanych aplikacji,
  • prowadzimy dokumentację oraz współpracujesz z dostawcami oprogramowania w zakresie obsługi błędów i problemów eksploatacyjnych,
  • bierzemy udział w pracach zespołów projektowych wdrażających usługi IT,
  • współpracujemy z administratorami zasobów informatycznych w zakresie eksploatowanych aplikacji,
  • testujemy i wdrażamy poprawki serwisowe i nowe wersje aplikacji,
  • bierzemy udział w szkoleniach i rozwoju w obszarze zarządzania tożsamością.

To Stanowisko Może Być Twoje, Jeśli

  • masz wiedzę z zakresu dostępnych narzędzi IDM na rynku, najlepiej znajomość OneIdentity Management,
  • masz wiedzę z zakresu zarządzania tożsamością użytkownika, proces J/M/L- doświadczenie w pracy z serwerami działającymi pod kontrolą systemów operacyjnych Microsoft Serwer, Linux (RHEL),
  • posiadasz doświadczenie z narzędziami z obszaru komponentów aplikacyjnych i sieciowych, takimi jak IIS, Windows Services, Certyfikaty SHA, Apache Tomcat, WebServices,
  • masz doświadczenie z bazami danych MS SQL\PostgreSQL w tym umiejętność pisania skryptów SQL, PLSQL,
  • masz doświadczenie w pracy z powłoką shell systemów Linux, pisanie skryptów, edytory,
  • masz doświadczenie w pracy z PowerShell w tym pisanie skryptów,
  • posiadasz umiejętność analizy logów systemowo/aplikacyjnych i na tej podstawie działań proaktywnych w celu eliminacji błędów,
  • znasz język angielski w stopniu umożliwiającym swobodne posługiwanie się dokumentacją techniczną.

Mile Widziane

  • znajomość narzędzi CI/CD: Jenkins, Docker Compose, GITLab etc.
  • podstawowa wiedza na temat technologii chmurowych (Microsoft Azure, Google Cloud).

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

121 more openings in this category and country

Software Engineering Manager, Site Reliability Engineering, Gemini Enterprise Agent Platform, Google CloudGoogle · Poland

Apply on the employer's site