Skip to content
Site Reliability Engineer

Site Reliability Engineer

Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure

NVIDIA AI

Warsawfulltimesenior

Откликнуться на сайте работодателя

Описание вакансии

Job Requisition ID

JR2015529

Job Category

Engineering

Time Type

Full time

NVIDIA's Deep Learning Frameworks (DLFW) Infrastructure team is looking for a deeply technical Senior HPC Cluster Administrator to lead the design, deployment, and reliability of our large-scale GPU compute clusters. These systems run the most demanding deep learning training, inference, and high-performance computing workloads in the industry — from DGX/HGX platforms to ground-breaking Grace Blackwell systems. You will drive architectural decisions across compute, networking, and storage, and partner closely with software, research, and product teams to keep our infrastructure ahead of the workloads it supports.

What You'll Be Doing

  • Own the full lifecycle of GPU compute clusters — procurement, provisioning, configuration management, monitoring, and deprecation — across heterogeneous Linux environments (DGX, HGX, embedded systems)
  • Design and scale storage solutions (NFS, Lustre, WekaFS, or equivalent) with a clear roadmap for capacity and performance growth
  • Lead automation of infrastructure using modern IaC tools (Ansible, Terraform) and CI/CD pipelines (GitLab)
  • Manage and optimize job scheduling via Slurm, including fair-share policies, reservation management, and MIG/GPU partitioning strategies
  • Maintain and improve observability stacks (Prometheus, Grafana, DCGM) and drive proactive resolution of hardware and software incidents
  • Collaborate with ML engineers and software teams to tune cluster configuration for large-scale distributed training workloads
  • Evaluate and introduce new technologies — networking fabrics (InfiniBand, NVLink, EFA/RDMA), storage tiers, container runtimes — to improve performance and reliability
  • Mentor junior engineers and contribute to team-wide engineering standards

What We Need To See

  • BS/MS in CS, EE, CE, or equivalent hands-on experience
  • 5+ years of experience deploying and administering large-scale HPC or ML training clusters
  • Deep expertise in Linux systems administration at scale
  • Strong scripting and automation skills in Python and/or bash
  • Hands-on experience with Slurm (scheduling, accounting, cgroup configuration)
  • Proficiency with configuration management and IaC (Ansible required; Terraform a plus)
  • Experience with container technologies (Docker, Apptainer/Singularity, Kubernetes)
  • Solid understanding of high-speed networking (InfiniBand, RoCE, RDMA, EFA)
  • Experience with distributed/parallel filesystems and storage architecture
  • Ability to own problems end-to-end and communicate clearly with engineering and management stakeholders

Ways To Stand Out From The Crowd

  • Experience with NVIDIA GPU infrastructure tools (DCGM, nvidia-smi, MIG, NVSwitch diagnostics)
  • Familiarity with cluster management platforms (Colossus, Bright Cluster Manager, xCAT, or similar)
  • Experience supporting large-scale distributed deep learning workloads (PyTorch, JAX, Megatron)
  • Knowledge of BMC/IPMI/Redfish for out-of-band management and hardware lifecycle
  • Background in MLOps tooling or ML platform engineering

Join our team of world-class engineers and be part of the groundbreaking work we do at NVIDIA. We are committed to encouraging a collaborative and inclusive environment, where every team member has the opportunity to thrive and make a significant impact!

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. For Poland: The base salary range is 221,250 PLN - 383,500 PLN for Level 3, and 292,500 PLN - 507,000 PLN for Level 4.

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

Скоро на этой странице

Резюме под эту вакансию — и билет в розыгрыш

Мы разбираем объявление до настоящих требований и переписываем ваше резюме под него — вопросами, а не выдумкой: ни одна строка не появится без вашего подтверждения. Войдите, чтобы получить это первым, — и попасть в розыгрыш.

  • Резюме под конкретную вакансию, а не «универсальное»
  • Ответы хранятся: правится любой, а не весь разговор заново
  • Всё в аккаунте — открывается с любого устройства

Разыгрываем

Скидка на сопровождение

Победителей выбираем случайно среди заявок с подтверждённой почтой. Дата розыгрыша и полные правила — на странице розыгрыша.

Правила розыгрыша

Site Reliability Engineer

Senior Cloud Engineer (Blockchain) - Digital Assets

EPAM Systems

Krakowcontractsenior

Откликнуться на сайте работодателя

Описание вакансии

We are looking for a
Senior Cloud Engineer
to join our Digital Assets Crew, working on pioneering blockchain-based projects including tokenization of financial assets. Our team operates within two Scrum teams and builds cloud-hosted applications that leverage cutting-edge technologies to deliver value and innovation to the finance industry

Responsibilities

  • Develop cloud solutions that scale and ensure high availability
  • Nurture a DevOps culture across the team
  • Contribute to DLT-related projects, particularly blockchain
  • Design and maintain infrastructure using Terraform and Helm
  • Implement GitOps and CI/CD pipelines
  • Monitor distributed systems using Grafana and Prometheus
  • Manage secret management solutions across cloud environments
  • Collaborate with colleagues in English on cross-functional initiatives

Requirements

  • 5+ years of experience in DevOps/SRE roles
  • Expertise in Azure Cloud and Kubernetes
  • Proficiency in Terraform, Helm and GitOps/CI-CD
  • Experience with Grafana, Prometheus and distributed systems architecture at enterprise scale
  • Knowledge of secret management solutions
  • Good communication skills, comfortable interacting with colleagues in English (B2+)

Nice to have

  • Experience with Istio Service Mesh and zero-trust architecture
  • Background in banking or regulated environments
  • Familiarity with blockchain technology

We offer

  • We gather like-minded people:
  • Top tech minds driving innovation in AI, cloud and digital platform modernization
  • Supportive team and agile, startup-like culture
  • Hybrid by design mode and opportunity to work remotely within Poland
  • Chance to work abroad for up to 60 days annually
  • Business-driven relocation opportunities
  • We provide growth opportunities:
  • Career development programs
  • Thought leadership, mentoring, soft skills and well-being programs
  • Certification (Anthropic, Gemini, GCP, Azure, AWS)
  • English classes
  • We cover it all:
  • Stable pay
  • Participation in the Employee Stock Purchase Plan with a 15% discount
  • Benefits package (health insurance, multisport, shopping vouchers)
  • Referral bonuses up to $2,000
  • Offices featuring entertainment and relaxation zones, table tennis and football, free snacks, coffee and more
  • Corporate, social and well-being events
  • Please, note:
  • Benefits listed above are available to employees only
  • We are open for working with Contractors. Terms of B2B cooperation agreements are agreed individually
  • We will reach out to selected candidates exclusively

EPAM is global leader in AI transformation engineering and integrated consulting, serving Forbes Global 2000 companies and ambitious startups. With over thirty years of expertise in custom software, product and platform engineering, we empower our clients to become AI-Native enterprises, driving measurable value from innovation and digital investments.

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

Site Reliability Engineer

ITIL Service Management Expert (EU Institution) - hybrid in Poland

NRB

Warsaw

Откликнуться на сайте работодателя

Описание вакансии

Job Description
Who are we?
KEYES
is a dynamic global organization that takes pride in being the trusted partner of
EU Institutions.
With strong commitment to excellence and a
30-years track record
of delivering high-quality solutions, we are dedicated to supporting the growth and success of our clients. Our Mission is to help our clients keep up with the challenges of
digital transformation
by providing the right talent at the right time for the right job. To this end, we are constantly looking for talented professionals who are interested in working on
challenging international projects
and able to deliver high-quality results within multicultural environments. Our services include (but are not limited to)
modernization of solutions, digital workspaces, cloud technologies and IT security
. Our Headquarters are in Brussels and we have active accounts and offices across Europe (i.e. Luxembourg, Amsterdam, Athens, Stockholm, Geneva).

Is this YOU?
For our customer based in
Warsaw, Poland -
an European Institution, we are looking for
ITIL Service Manager Expert based in Poland
to join a long-term mission in the area of
cybersecurity, public sector
and
law enforcement
. You will play a key role in managing business architecture, guiding solution designs, and fostering collaboration among business architects and project teams to ensure coherence with the customer’s long-term vision.

Please note that the role requires 80-100% hybrid presence at the client’s headquarters in Warsaw, meaning that being based 2-3 hours by ground transportation from Warsaw will be required.
More specifically, you will be responsible for…

  • Supporting execution of the existing IT service management processes
  • Analysing existing IT service management processes
  • Identifying risks and weak points and recommending improvements
  • Advising on the best IT service management practices
  • Preparing, documenting and implementing improvement plans
  • Supporting implementation and customisation of ITSM solution
  • Participation in projects related with implementation and customization of ITSM solution
  • Preparing required policies and Standard Operating Procedures (SOPs) supporting IT service management processes
  • Creating a positive customer experience
  • Other specific duties as assigned by the ICT SCSMT team leader.

Job Requirements
Are you the perfect match?

  • University degree (BSc/MSc)
  • At least 3 years (full time) experience at the similar position
  • At least 2 years (full time) experience in managing ICT service delivery and/or ICT operations
  • At least ITIL Intermediate certificate (ITIL Expert certificate preferred)
  • Excellent practical knowledge of IT Service Management
  • Practical experience with ITSM solutions
  • Well understanding of complex information systems and their interoperability
  • Excellent understanding and practical knowledge of IT technologies and information systems technical components
  • Familiarity with project management approaches, tools and phases of the project lifecycle Skills:
  • Excellent communication skills (in written and verbal communication)
  • Very strong sense of responsibility
  • Accuracy and attention to details
  • Very good organizational skills
  • Very good reporting skills
  • Supportive and helpful personality with co-operative and service oriented attitude
  • Forward looking with a holistic approach
  • High level of motivation and initiative.

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

Site Reliability Engineer

IT/OT Systems Engineer

ALGOTEQUE Innovation Hub

Warsawsenior

Откликнуться на сайте работодателя

Описание вакансии

ALGOTEQUE is an IT consultancy firm that helps startups, mid-sized and large corporations to create and deliver innovative technologies.

Our team has a successful track record in designing, developing, implementing, and integrating software solutions (AI, ML, BI, Web, Automation) for Telecom, Energy, Bank, Insurance, Pharma, Automotive, Industry, e-commerce. We deliver our services both in fixed-price and time-and-materials models, helping our customers achieve their business and IT strategies.

Job Description
Poszukujemy Inżyniera Systemów IT/OT, który wesprze rozwój i utrzymanie infrastruktury technologicznej zakładu produkcyjnego. Konsultant będzie kluczowym ogniwem łączącym świat IT (infrastruktura, sieci, systemy) z OT (automatyka, linie produkcyjne, systemy sterowania), wspierając inicjatywy z obszaru Industry 4.0.

Zakres Projektu

  • Utrzymanie ciągłości działania systemów produkcyjnych (SCADA, HMI, MES)
  • Integracja systemów OT z infrastrukturą IT (serwery, sieci, środowiska chmurowe)
  • Wsparcie cyfryzacji procesów produkcyjnych i wdrożeń Industry 4.0
  • Diagnostyka i rozwiązywanie problemów na styku IT/OT (incydenty, wydajność)
  • Zarządzanie infrastrukturą sieciową w środowisku produkcyjnym (VLAN, segmentacja)
  • Współpraca z działami utrzymania ruchu, automatykami i zespołami IT
  • Wdrażanie i utrzymanie standardów bezpieczeństwa w środowisku OT
  • Tworzenie dokumentacji technicznej oraz procedur operacyjnych

Profile / Requirements

  • 3–5 lat doświadczenia w środowiskach IT/OT lub infrastrukturalnych
  • Praktyczna znajomość systemów przemysłowych (SCADA, MES, PLC – mile widziane)
  • Doświadczenie w pracy w środowisku produkcyjnym (fabryka / zakład przemysłowy)
  • Znajomość sieci przemysłowych i IT (TCP/IP, VLAN, routing, firewall)
  • Systemy operacyjne: Windows Server i/lub Linux
  • Umiejętność pracy w środowisku krytycznym (wysoka dostępność, uptime)
  • Język angielski min. B2

Mile Widziane

  • Znajomość standardów bezpieczeństwa OT (ISA/IEC 62443)
  • Doświadczenie z systemami MES / integracją danych produkcyjnych
  • Chmura (Azure / AWS) w kontekście zbierania i analizy danych z produkcji
  • Podstawy automatyki (PLC – Siemens, Rockwell)
  • Skrypty / automatyzacja (PowerShell, Python)

AO4344

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

Site Reliability Engineer

Administratorka/Administrator aplikacji

PKO Bank Polski

Warsaw

Откликнуться на сайте работодателя

Описание вакансии

Na Co Dzień w Naszym Zespole

  • administrujemy aplikacjami z obszarów zarządzania tożsamością, HR oraz gospodarki własnej Banku,
  • utrzymujemy dostępność aplikacji administrowanych przez nasz zespół poprzez ciągłe monitorowanie zasobów jak i procesów, wgrywanie zmian oraz poprawek aplikacyjnych,
  • realizujemy usługi informatyczne dla jednostek biznesowych banku w zakresie eksploatowanych aplikacji,
  • rozwiązujemy incydenty/problemy zgłaszane przez jednostki banku z zakresu działania administrowanych aplikacji,
  • prowadzimy dokumentację oraz współpracujesz z dostawcami oprogramowania w zakresie obsługi błędów i problemów eksploatacyjnych,
  • bierzemy udział w pracach zespołów projektowych wdrażających usługi IT,
  • współpracujemy z administratorami zasobów informatycznych w zakresie eksploatowanych aplikacji,
  • testujemy i wdrażamy poprawki serwisowe i nowe wersje aplikacji,
  • bierzemy udział w szkoleniach i rozwoju w obszarze zarządzania tożsamością.

To Stanowisko Może Być Twoje, Jeśli

  • masz wiedzę z zakresu dostępnych narzędzi IDM na rynku, najlepiej znajomość OneIdentity Management,
  • masz wiedzę z zakresu zarządzania tożsamością użytkownika, proces J/M/L- doświadczenie w pracy z serwerami działającymi pod kontrolą systemów operacyjnych Microsoft Serwer, Linux (RHEL),
  • posiadasz doświadczenie z narzędziami z obszaru komponentów aplikacyjnych i sieciowych, takimi jak IIS, Windows Services, Certyfikaty SHA, Apache Tomcat, WebServices,
  • masz doświadczenie z bazami danych MS SQL\PostgreSQL w tym umiejętność pisania skryptów SQL, PLSQL,
  • masz doświadczenie w pracy z powłoką shell systemów Linux, pisanie skryptów, edytory,
  • masz doświadczenie w pracy z PowerShell w tym pisanie skryptów,
  • posiadasz umiejętność analizy logów systemowo/aplikacyjnych i na tej podstawie działań proaktywnych w celu eliminacji błędów,
  • znasz język angielski w stopniu umożliwiającym swobodne posługiwanie się dokumentacją techniczną.

Mile Widziane

  • znajomość narzędzi CI/CD: Jenkins, Docker Compose, GITLab etc.
  • podstawowa wiedza na temat technologii chmurowych (Microsoft Azure, Google Cloud).

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

Ещё 120 вакансий по этой категории в этой стране

Senior HPC Cluster Administrator - Deep Learning Frameworks InfrastructureNVIDIA AI · Poland

Откликнуться на сайте работодателя
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure — NVIDIA AI | mentors.coach