Skip to content
Site Reliability Engineer

Site Reliability Engineer

Lead AI Engineer, Agentic and RAG Systems

EPAM Systems

Удалённо

Откликнуться на сайте работодателя

Описание вакансии

We are looking for a seasoned
Lead AI Engineer
who architects, builds, and operates production GenAI platforms – agentic workflows, RAG pipelines, and LLM-backed services with real users and real SLAs – while leading engineers and setting the technical direction across multiple workstreams.

This is an engineering leadership role, not a research role. The bar is reliability, latency, cost, observability, and safe deployment at scale, with end-to-end ownership from architecture through on-call, and accountability for the technical quality and delivery of the team. Typical workloads include enterprise knowledge platforms, conversational analytics, agentic automation, and LLM-augmented data products.

This is a fully remote position that offers you the flexibility to work from any location in Armenia, whether it's your home or well-equipped offices in Yerevan or Gyumri.
Responsibilities

  • Own the end-to-end architecture of GenAI platforms across multiple services and teams, defining standards, patterns, and reference implementations
  • Lead the design of agent orchestration (graph/state, conditional routing, tool calling, memory, checkpointing) in LangGraph / LangChain or equivalent, and set best practices for the team
  • Architect production RAG end-to-end: chunking, embeddings, vector stores, hybrid retrieval, reranking, caching, and grounded synthesis – and mentor engineers in building it
  • Drive the design and delivery of Python / FastAPI services – async, SSE streaming, session handling, and structured error contracts – establishing service templates and conventions
  • Define the observability and evaluation strategy (MLflow, OpenTelemetry, or equivalent) for accuracy, cost, and regression across the platform
  • Own the deployment platform on Docker + Kubernetes (EKS/AKS/GKE) with CI/CD, test, eval, and canary gates – setting release standards for AI systems
  • Lead LLM cost engineering strategy – model routing, prompt optimization, caching, token accounting, and build-vs-buy decisions at portfolio level
  • Establish GenAI safety & governance practices: hallucination control, prompt-injection defense, PII handling, and HITL where required
  • Partner with data engineering leadership on semantic layers and pipelines (PySpark / SQL where applicable), and align roadmaps across teams
  • Mentor and grow senior and mid-level engineers through design reviews, pairing, and technical coaching; conduct hiring and technical interviews
  • Represent engineering in conversations with clients, product, and executive stakeholders; translate business goals into technical strategy and delivery plans

Requirements

  • 6+ years in software engineering, with 3+ years shipping production LLM / agentic systems (not POCs or research)
  • 1+ years of experience leading engineers or technical workstreams
  • Proven track record of owning architecture for multi-service GenAI or distributed systems in production
  • Expert-level proficiency in Python and FastAPI (async, REST, SSE)
  • Deep production expertise in LangChain and LangGraph (or equivalent serious production experience with LlamaIndex, AutoGen, or MCP stacks)
  • Strong background in production RAG: embeddings, chunking, and hybrid retrieval with reranking and caching – with the ability to define standards across teams
  • Advanced skills in vector databases such as Pinecone, Weaviate, pgvector, OpenSearch, or Databricks Vector Search
  • Hands-on production experience with at least one major LLM provider – AWS Bedrock (preferred), OpenAI / Azure OpenAI, or Anthropic – including model selection, routing trade-offs, and multi-provider strategy
  • Strong competency in Kubernetes and Docker in real production environments (EKS/AKS/GKE), including platform-level decisions
  • Deep expertise in cloud engineering on AWS, including cost, security, and scalability trade-offs
  • Solid command of observability and tracing tools (MLflow, LangSmith, OpenTelemetry), evaluation harnesses, and latency/cost ownership at platform scale
  • Experience designing and owning CI/CD for AI systems (GitHub Actions, Jenkins, or equivalent) with test/eval gates
  • Demonstrated experience mentoring engineers, leading design reviews, and driving technical decisions across teams
  • Strong written and spoken English (B2+ level); able to lead design discussions, present to senior stakeholders, and influence technical direction with clients and executives

Nice to have

  • Databricks depth – MLflow (tracking & serving), Vector Search, Unity Catalog / Metric Views, PySpark / SQL
  • Experience with LLM fine-tuning – PEFT, LoRA, QLoRA – and the ability to guide build-vs-fine-tune-vs-prompt decisions
  • Strong understanding of MCP servers and tool integration patterns
  • Expertise in GenAI governance & FinOps – auditability, prompt-injection hardening, PII, and token cost in regulated environments
  • Background in classical ML / DL – NLP, BERT-family, time-series, and CV

We offer

  • We connect like-minded people
  • Delivering innovative solutions to industry leaders, making a global impact
  • Enjoyable working environment, whether it is the vibrant office or the comfort of your home
  • Opportunity to work abroad for up to two months per year
  • Relocation opportunities within our offices in 55+ countries
  • Corporate and social events
  • We invest in your growth
  • Leadership development, career advising, soft skills and well-being programs
  • Certifications, including GCP, Azure and AWS
  • Unlimited access to EPAM's internal learning database
  • Free English classes with certified teachers
  • We cover it all
  • Participation in the Employee Stock Purchase Plan
  • Monetary bonuses for engaging in the referral program
  • Comprehensive medical & family care package
  • Four trust days per year for personal needs
  • Discounts for fitness clubs
  • Benefits package (hotels, restaurants, stores and services)

EPAM is global leader in AI transformation engineering and integrated consulting, serving Forbes Global 2000 companies and ambitious startups. With over thirty years of expertise in custom software, product and platform engineering, we empower our clients to become AI-Native enterprises, driving measurable value from innovation and digital investments.

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

Скоро на этой странице

Резюме под эту вакансию — и билет в розыгрыш

Мы разбираем объявление до настоящих требований и переписываем ваше резюме под него — вопросами, а не выдумкой: ни одна строка не появится без вашего подтверждения. Войдите, чтобы получить это первым, — и попасть в розыгрыш.

  • Резюме под конкретную вакансию, а не «универсальное»
  • Ответы хранятся: правится любой, а не весь разговор заново
  • Всё в аккаунте — открывается с любого устройства

Разыгрываем

Скидка на сопровождение

Победителей выбираем случайно среди заявок с подтверждённой почтой. Дата розыгрыша и полные правила — на странице розыгрыша.

Правила розыгрыша

Site Reliability Engineer

Ingénieur Système, Sauvegarde et réseaux à Paris H/F

Free-Work

Удалённоfulltimesenior

Откликнуться на сайте работодателя

Описание вакансии

Dans le cadre du développement de notre équipe IT chez l'un de nos clients grands comptes, nous recherchons un(e) Ingénieur(e) Système Sauvegarde réseaux H/F afin d'assurer l'administration, la maintenance et l'évolution des infrastructures systèmes de nos environnements clients.

Vos missions seront les suivantes :

  • Participer aux projets d'évolutions de la plateforme technique de la Video Factory ( tête de réseau OTT et IPTV
  • Conception / participation au POC avec le N3 (les experts) / intégration / ingénierie / recette unitaire / recette des workflow et documentation des briques techniques et des procédures.
  • Assurer la maintenance en condition opérationnelle de la plateforme
  • Organisation des sauvegardes et des montées de version des équipements IT (VM, NAS, OS (linux et windows)...) et Vidéos (encodeurs, DCM, serveurs d'origine, sondes ) de la plateforme.
  • Analyse des risques et des impacts potentiels et planification en HNO le cas échéant.
  • Assurer le « run » de la plateforme en tant que support niveau 2 en soutien des équipes support de niveau 0 et 1
  • Analyse d'incidents et résolutions, escalade au N3 et aux fournisseurs (ouverture et suivi des tickets) le cas échéant, communication sur les avancées les plus significatives et les impacts majeurs,
  • Organiser les activités des fournisseurs (mises à jour et évolution technique),
  • Assurer le suivi des déploiements et mettre en place les contrôles (recette).
  • Participer à l'amélioration de l'organisation du support global de la plate-forme technique :
  • Formation des équipes de maintenance, documentation des procédures d'exploitation.

     Des missions complémentaires peuvent être confiées.

Référence de l'offre : cmk6d3cswx

Profil candidat:

Profil recherché :

Issu(e) d'une formation supérieure en informatique (BAC+5 / Diplôme d'ingénieur),

Vous justifiez d'au moins 5 ans d'expérience sur un poste similaire.

Compétences techniques souhaitées :

  • Maîtrise des environnements Windows Server, Linux et Vmware (des connaissances sur Nutanix est un plus).
  • Bonne connaissance des environnements vidéo et audio sur IP.
  • Connaissance des réseaux IP et de l'adressage multicast.
  • Maitrise des outils d'analyse de qualité vidéo ainsi que des outils de supervision.
  • Bonnes capacités d'analyse et de résolution de problèmes.
  • Autonomie, rigueur et bon relationnel.
  • Capacité à travailler en équipe et à intervenir dans des environnements de production

Environnement technique :

DCM, Anevia, Imagine, Nevion, Harmonic xOS, OpenHeadEnd, Elemental, USP.

Produits systèmes, virtualisation : VMware, Wallix, Nutanix, NAS, Debian, Windows Server, FTP

Ce que nous vous proposons :

Valeurs : en plus de nos 3 fondamentaux que sont l'audace, la bonne foi et la réactivité, nous garantissons un management à l'écoute et de proximité, ainsi qu'une ambiance familiale.

Contrat : CDI ou Freelance

Localisation : Paris

Package rémunération & avantages :

  • Le salaire : rémunération annuelle brute selon profil et compétences
  • Les basiques : mutuelle familiale et prévoyance, titres restaurant, remboursement transport en commun à 50%, avantages du CSE (culture, voyage, chèque vacances, et cadeaux), RTT (jusqu'à 12 par an), plan d'épargne, prime de participation
  • Nos plus : forfait mobilité douce & Green (vélo/trottinette et covoiturage), prime de cooptation de 1000 € brut, e-shop de matériel informatique à des prix préférentiels (smartphone, tablette, etc.)
  • Votre carrière : plan de carrière, dispositifs de formation techniques & fonctionnels, passage de certifications, accès illimité à Microsoft Learn
  • La qualité de vie au travail : télétravail avec indemnité, évènements festifs et collaboratifs, accompagnement handicap et santé au travail, engagements RSE

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

Site Reliability Engineer

SAP DevOps Engineer (f/m/d)

E.ON Digital Technology

Удалённо

Откликнуться на сайте работодателя

Описание вакансии

You have a passion for technology and want to make the world a greener place?
Then become a playmaker (f/m/d) and join our team as SAP DevOps Engineer (f/m/d) at E.ON Digital Technology.

We play a key role in shaping the energy transition by leading E.ON's digital transformation across Europe. We explore new paths by developing ideas, breaking new ground, making visions reality, and bringing new technologies to life. We deliver sustainable technology solutions because…

… it’s on us to make new energy work!
The Team
– your impact

At E.ON, the SAP Engineering Chapter is a collective of innovative minds dedicated to delivering world-class SAP architecture, SAP software development and SAP engineering capabilities across our segments and product teams. By joining us, you will play a critical role in keeping E.ON's SAP architecture and capabilities modern, secure and excellent.

Your Role –
meaningful & rewarding

As a SAP DevOps Engineer (f/m/d) at E.ON, you will be responsible for designing, developing, training and maintaining SAP DevOps solutions for our SAP system landscape and platforms. You will work closely with the SAP DevOps teams and developers, and other stakeholders to ensure the provisioning of a modern, user-friendly, state-of-the-art SAP DevOps pipeline.

  • Design, implement and continuously improve DevOps concepts, CI/CD templates, pipelines and automation for SAP landscapes (e.g. SAP BTP, S/4HANA)
  • Enable reliable build, test and deployment processes across SAP Cloud and on-premise environments
  • Establish observability concepts for SAP BTP and on-premise applications for monitoring, health checks, alerts and APM by using tools such as SAP Cloud ALM, New Relic and Uptrends
  • Collaborate closely with SAP development, architecture, security and operations teams
  • Conduct DevOps maturity assessments and guide teams on their DevOps Journey
  • Ensure high availability, performance, security and compliance of SAP systems
  • Troubleshoot complex problems and support root cause analysis
  • Continuously evaluate new SAP DevOps tools, technologies and best practices

Your Profile
– authentic & open-minded

  • Strong experience in DevOps, platform engineering or operations within SAP environments
  • Hands-on experience with CI/CD and security, as well as quality tooling (e.g. GitLab CI/CD, SonarQube, Renovate or similar)
  • Strong knowledge in development, automated testing and testing strategies
  • Experience with scripting and automation (e.g. Bash, Groovy)
  • Solid knowledge of SAP S/4HANA, SAP BTP and related deployment and transport mechanisms
  • Knowledge of Kubernetes and container technologies and IaaC principles is a plus
  • Strong understanding of security, monitoring and reliability concepts
  • Structured, proactive and solution-oriented working style
  • Fluent in English

Our Benefits
– smart & useful

  • Advance your development: We grow and we want you to grow with us. Learning on the job, exchanging with others, or taking part in an individualtraining – our learning culture enables you to bring your personal and professional development to the next level.
  • Recharge your battery: You have 30 days of paid vacation per year plus Christmas and New Year's Eve off. Your battery still needs charging? You canexchange parts of your salary for more paid vacation or you can take a sabbatical.
  • Enjoy hybrid work: We combine office collaboration with focused work from home. It’s also possible to go on workation for up to 20 days per year withinEurope.
  • Stay active & healthy: Benefit from a company-sponsored health membership.
  • Elevate your mobility: From car and bike leasing offers to a subsidised Deutschland-Ticket – your way is our way.
  • Think ahead: With our company pension scheme and a great insurance package we take care of your future.
  • This is by far not all… We are looking forward to speaking with you about further benefits during the hiring process.

Do you have questions?
For further information please contact Agneta Lierl, EDT_Talent_Acquisition@eon.com.

What you need to know:
Contract type: Permanent

Working time: Full time

Company: E.ON Digital Technology GmbH

Location: Essen, Hannover, München, Berlin, Würzburg, Hamburg, Frankfurt am Main

Function area: IT/Digital; Engineering

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

Site Reliability Engineer

Senior Security Engineer (m/w/d)

Rocken®

Удалённоsenior

Откликнуться на сайте работодателя

Описание вакансии

Steuerverwaltungen brauchen clevere Software – und dahinter stehen clevere Köpfe. Unser Rocken Partner entwickelt die führende Business Lösung für kantonale und kommunale Steuerverwaltungen. Die Anwendung deckt den kompletten Verwaltungsprozess ab: vom Steuerregister über Veranlagungen und Fakturierung bis zum Inkasso und zur Verlustscheinbewirtschaftung. Derzeit entsteht eine neue Software-Generation – ein spannendes Projekt, das technisches Know-how und Innovationskraft vereint. Die Arbeitsweise ist geprägt von überlegter Fokussierung, lebendigem Austausch zwischen Teams und intellektueller Courage. Elegante Lösungen entstehen durch Wissen, Geist und Teamenergie. Bereit, an einem Produkt zu arbeiten, das echten Impact hat? Unser Rocken Partner sucht Menschen, die sich für eine gemeinsame Idee begeistern und die digitale Zukunft der Steuerverwaltung mitgestalten möchten.

Verantwortung

  • Du analysierst Schwachstellen und setzt passende Sicherheitslösungen um.
  • Du arbeitest an Themen wie Zugriff, Netzwerk, Pentesting und Monitoring.
  • Du entwickelst die Sicherheitsarchitektur weiter und beachtest ISO 27001.
  • Du automatisierst in Windows/Linux und unterstützt Kubernetes-Setups.

Qualifikationen

  • Du hast eine IT-Ausbildung und Erfahrung als System Engineer mit Security-Fokus.
  • Du denkst vernetzt, erkennst Risiken und arbeitest agil mit modernen Tools.
  • Du kennst Azure, VMware und gängige Security-Lösungen.
  • Du brennst für IT-Security und entwickelst im Team nachhaltige Lösungen.

Benefits

  • Interessante und abwechslungsreiche Tätigkeiten/Projekte
  • Attraktive Weiterbildungs- und Entwicklungsmöglichkeiten
  • Flexible Arbeitszeitgestaltung
  • Homeoffice
  • Offene Unternehmenskultur
  • Beteiligung oder Übernahme ÖV-Abonnements
  • Beteiligung oder Übernahme Parkplatz
  • Attraktive Mitarbeiterrabatte
  • Kostenlose Früchte und Getränke

ROCKEN Jobs
https://rocken.jobs

Profil Erstellen
https://rocken.jobs/application/profil\-erstellen/

  • new

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

Site Reliability Engineer

Microsoft Cloud & System Engineer (m/w/d)

Rocken®

Удалённо

Откликнуться на сайте работодателя

Описание вакансии

Unser Rocken® Partner ist spezialisiert auf IT-Security, Cloud Lösungen, sowie IT-Outsourcing und Support. Sie setzen auf innovative Services und Dienstleistungen, bewegen und orientieren sich am Puls der Technik. Daraus ergeben sich für unsere Kunden viele Vorteile gegenüber den traditionellen IT-Lösungen, die weniger flexibel und selten skalierbar sind.

Verantwortung

  • Sicherstellung einer stabilen IT-Infrastruktur und hohen Systemverfügbarkeit
  • Betreuung und Weiterentwicklung von Microsoft-Umgebungen im Kundenumfeld
  • Planung und Umsetzung von Infrastruktur-, Rollout- und Migrationsprojekten
  • Analyse und Behebung technischer Störungen im 1st- und 2nd-Level-Support
  • Direkter Kundensupport, technische Beratung und Betreuung vor Ort

Qualifikationen

  • Microsoft 365, Azure, Entra ID und Microsoft Intune
  • Kenntnisse in Client-/Server-Systemen, Virtualisierung, Monitoring und Backup
  • Erfahrung im 1st-/2nd-Level-Support und technischen Troubleshooting
  • Kenntnisse in Infrastrukturplanung, System-Rollouts und Migrationen
  • Sehr gute Deutschkenntnisse sowie gute Englischkenntnisse; weitere Sprachen von Vorteil

Benefits

  • Flexible Arbeitszeitgestaltung
  • Homeoffice
  • Zahlreiche Mitarbeiterevents
  • Beteiligung oder Übernahme Parkplatz
  • Kostenlose Früchte und Getränke
  • Attraktive Weiterbildungs- und Entwicklungsmöglichkeiten

ROCKEN Jobs
https://rocken.jobs

Profil Erstellen
https://rocken.jobs/application/profil\-erstellen/

  • new

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

Ещё 241 вакансия по этой категории в этой стране

Lead AI Engineer, Agentic and RAG SystemsEPAM Systems

Откликнуться на сайте работодателя
Lead AI Engineer, Agentic and RAG Systems — EPAM Systems | mentors.coach