Skip to content
SRE

SRE

Senior Azure Cloud Engineer

EPAM Systems

Remotesenior

Apply on the employer's site

Role description

We are seeking an experienced
Senior Azure Cloud Engineer
to design, build, automate, and operate secure, scalable, and cost-efficient cloud solutions on Microsoft Azure.

In this role, you will work hands-on with Azure infrastructure, containers, DevOps pipelines, Infrastructure as Code, observability, security controls, and AI-enabled cloud solutions. You will help translate architecture standards into working platforms, reusable deployment patterns, automation, and production-ready services.

You will collaborate closely with Cloud Architects, Platform Engineering, DevOps, Security, Networking, Data, AI, and Product teams to deliver robust Azure solutions that support modern application delivery, containerized workloads, and AI-enabled enterprise capabilities.

 

This is a fully remote position that offers you the flexibility to work from any location in Armenia, whether it's your home or well-equipped offices in Yerevan or Gyumri.

 

 

Responsibilities

  • Design, implement, and maintain Azure cloud infrastructure, including subscriptions, resource groups, networking, identity, governance, security, and platform services
  • Build and support reusable Azure deployment patterns for applications, APIs, front-end workloads, microservices, containers, serverless services, data integrations, and AI-enabled solutions
  • Implement and operate Azure Kubernetes Service environments, including node pools, ingress, workload identity, secrets management, autoscaling, networking, monitoring, container registry integration, and security controls
  • Build, maintain, and improve Infrastructure as Code using Terraform, including reusable modules, multi-environment deployments, automated validation, and integration with CI/CD pipelines
  • Design and implement CI/CD pipelines using GitHub Actions, Azure DevOps, GitLab CI/CD, or similar tools
  • Support DevOps practices such as automated builds, testing, security scanning, artifact management, environment promotion, deployment approvals, rollback strategies, and release automation
  • Use GitHub Copilot and AI-assisted engineering tools to improve productivity across scripting, IaC development, CI/CD pipeline creation, code review, troubleshooting, documentation, and automation
  • Implement cloud-native solutions using Azure services such as App Service, Azure Functions, Logic Apps, Event Grid, Service Bus, Storage, Key Vault, API Management, Azure SQL, Cosmos DB, Azure Monitor, and Application Insights
  • Support the implementation of AI-enabled solutions using Azure AI, Azure OpenAI, Azure AI Search, Azure Machine Learning, and related Azure AI services
  • Build and integrate AI solution components based on patterns such as RAG, agentic workflows, multi-agent orchestration, tool/function calling, prompt management, grounding, evaluation, and responsible AI controls
  • Implement secure integration patterns using managed identities, RBAC, private endpoints, private DNS, network security groups, firewalls, and API gateways
  • Establish and maintain observability using Azure Monitor, Log Analytics, Application Insights, Container Insights, dashboards, alerts, distributed tracing, and operational runbooks
  • Troubleshoot complex cloud, networking, deployment, performance, security, and production incidents
  • Apply DevSecOps practices, including secrets management, dependency scanning, container image scanning, policy validation, secure configuration, and compliance automation
  • Optimize Azure environments for performance, reliability, scalability, and cost efficiency
  • Contribute to technical standards, reusable templates, documentation, operational procedures, and platform engineering practices
  • Provide technical guidance, mentoring, code reviews, and engineering leadership to other team members
  • Work with architects and stakeholders to translate requirements into practical, secure, and maintainable Azure implementations

Requirements

  • Strong hands-on experience designing, implementing, and operating Azure cloud solutions in enterprise environments
  • Solid understanding of Azure networking, identity, governance, security, monitoring, and platform services
  • Practical experience with Azure Landing Zone concepts, hub-and-spoke networking, private endpoints, private DNS, firewalls, NSGs, route tables, and workload integration patterns
  • Strong hands-on experience with Azure Kubernetes Service, containers, container registries, ingress controllers, workload identity, autoscaling, monitoring, and container security
  • Strong experience with Terraform or other Infrastructure as Code tools, including module development, state management, validation, and multi-environment delivery
  • Strong experience with CI/CD pipelines, preferably using GitHub Actions, Azure DevOps, GitLab CI/CD, or similar platforms
  • Good understanding of DevOps and DevSecOps practices, including automated testing, security scanning, artifact management, release automation, and deployment governance
  • Experience working with GitHub, pull requests, code reviews, branching strategies, and collaborative engineering workflows
  • Practical experience using GitHub Copilot or similar AI-assisted development tools for infrastructure, automation, scripting, pipeline development, or documentation
  • Experience with Azure PaaS and integration services such as App Service, Azure Functions, Logic Apps, API Management, Event Grid, Service Bus, Storage, Key Vault, Azure SQL, Cosmos DB, and related services
  • Experience implementing observability using Azure Monitor, Log Analytics, Application Insights, Container Insights, dashboards, alerting, and operational runbooks
  • Understanding of Azure AI and Generative AI services, especially Azure OpenAI, Azure AI Search, and AI-enabled automation patterns
  • Practical knowledge of AI architecture patterns such as RAG, agentic workflows, multi-agent systems, tool/function calling, grounding, prompt management, evaluation, and responsible AI
  • Ability to troubleshoot complex technical issues across cloud infrastructure, networking, containers, CI/CD, security, and application integration
  • Strong scripting and automation skills using PowerShell, Bash, Python, or similar languages
  • Ability to work independently, take ownership of technical delivery, and support production-grade cloud environments
  • Strong communication skills and ability to work with architects, engineers, security teams, product teams, and business stakeholders

Nice to have

  • Microsoft Azure certifications such as Azure Administrator Associate, Azure Developer Associate, Azure DevOps Engineer Expert, or Azure Solutions Architect Expert
  • Kubernetes certifications such as CKA, CKAD, or CKS
  • Experience with production-grade AKS platforms, service mesh, GitOps, Helm, Kustomize, Flux, Argo CD, or Kubernetes policy engines
  • Experience with Azure AI Foundry, Semantic Kernel, LangChain, LangGraph, AutoGen, or similar AI orchestration frameworks
  • Experience building or supporting RAG platforms, AI agents, multi-agent workflows, or enterprise knowledge search solutions
  • Experience with platform engineering, internal developer platforms, self-service cloud capabilities, paved roads, and reusable engineering templates
  • Experience working in regulated industries with strong compliance, security, auditability, and governance requirements
  • Familiarity with SRE practices, incident response, reliability engineering, performance testing, and cost optimization

 

We offer

  • We connect like-minded people
    • Delivering innovative solutions to industry leaders, making a global impact
    • Enjoyable working environment, whether it is the vibrant office or the comfort of your home
    • Opportunity to work abroad for up to two months per year
    • Relocation opportunities within our offices in 55+ countries
    • Corporate and social events
  • We invest in your growth
    • Leadership development, career advising, soft skills and well-being programs
    • Certifications, including GCP, Azure and AWS
    • Unlimited access to EPAM's internal learning database
    • Free English classes with certified teachers
  • We cover it all
    • Participation in the Employee Stock Purchase Plan
    • Monetary bonuses for engaging in the referral program
    • Comprehensive medical & family care package
    • Four trust days per year for personal needs
    • Discounts for fitness clubs
    • Benefits package (hotels, restaurants, stores and services)

 

EPAM is global leader in AI transformation engineering and integrated consulting, serving Forbes Global 2000 companies and ambitious startups. With over thirty years of expertise in custom software, product and platform engineering, we empower our clients to become AI-Native enterprises, driving measurable value from innovation and digital investments.

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Coming to this page

A resume for this role — and a ticket to the draw

We take the posting apart down to the real requirements and rewrite your resume against it — by asking, not inventing: no line appears without your confirmation. Sign in to get it first, and to enter the draw.

  • A resume for this exact role, not a universal one
  • Answers are kept: edit any one, not the whole conversation
  • All in your account — open it from any device

On the wheel

A discount on mentoring

Winners are drawn at random among entries with a confirmed email. The date and the full rules are on the draw page.

Draw rules

SRE

Expert systèmes ( F/H )

Plaine Commune

Remote

Apply on the employer's site

Role description

A proximité immédiate de Paris, Plaine Commune est un Territoire de 450 000 habitants, composé des villes d’Aubervilliers, Epinay-sur-Seine, La Courneuve, L’Ile-Saint-Denis, Pierrefitte-sur-Seine, Saint Denis, Saint- Ouen, Stains et Villetaneuse. Elles sont fédérées autour d’un projet commun, sur un espace qui connaît des mutations inédites en région parisienne. Plaine Commune exerce des activités essentielles comme l'aménagement urbain, le développement économique et les services à la population (gestion de l'espace public : propreté, espaces verts, Lecture publique...).

La direction des systèmes d’information mutualisés de Plaine Commune pilote les systèmes d’information, la gestion administrative et financière, ainsi qu’un portefeuille de projets informatiques pour le compte de quatre entités : l’établissement public Plaine Commune, la Ville de Saint-Denis, la Ville de Villetaneuse et la Ville de L’Île-Saint-Denis.

L’Expert Systèmes assure la disponibilité, la performance et la sécurité des systèmes informatiques en production. Il définit les standards techniques, pilote les projets d’infrastructure et veille à leur mise en œuvre, dans une logique d’amélioration continue et de fiabilité du système d’information

Rattachement hiérarchique du poste
: Responsable de l’équipe Moyens Technologiques

Missions principales

Administration et gestion des systèmes

  • Garantir la cohérence, la fiabilité et la performance des infrastructures systèmes (serveurs, stockage, sauvegardes, supervision).
  • Assurer la maintenance opérationnelle et la continuité de service.
  • Gérer les changements en production et suivre la qualité du service.
  • Encadrer les aspects administratifs : gestion fournisseurs, bons de commande, marchés publics.

Mise en œuvre de projets techniques

  • Piloter les projets d’infrastructure et de sécurité (serveurs, antivirus, messagerie, sauvegardes, etc.).
  • Définir et tester les architectures techniques.
  • Encadrer les prestataires pour garantir performance et respect contractuel.
  • Contribuer aux projets transverses du SI.

Sécurité et veille technologique

  • Assurer la conformité et la sécurité des systèmes (paramétrages, MFA, DLP, Defender, politiques d’accès).
  • Surveiller et maintenir la résilience face aux menaces.
  • Réaliser une veille technologique active et proposer des améliorations.

Gestion des référentiels utilisateurs

  • Administrer les comptes et services liés aux référentiels techniques (Active Directory, messagerie, téléphonie).
  • Gérer les environnements Microsoft 365 (licences, configurations, OneDrive, SharePoint, Teams).
  • Garantir la conformité des accès et la collaboration sécurisée.

Activités occasionnelles

  • Contribuer la continuité du service en l’absence de collègues et/ou du-de la Responsable
  • Participer à la cellule de crise en cas d’incident grave sur le Système d’Information

Savoirs

  • Architecture systèmes et sécurité informatique.
  • Maîtrise des infrastructures et procédures de maintien en condition opérationnelle.

Savoir-faire

  • Expertise sur environnements Windows, Linux, VMWare, Microsoft 365, et sécurité Cloud.
  • Solides compétences en gestion de projet et en supervision technique.
  • Capacité à gérer plusieurs tâches simultanément et à intervenir en urgence.
  • Expérience dans la gestion de prestataires et marchés publics.

Savoir-être

  • Sens du service public et de la qualité de service.
  • Autonomie, réactivité et disponibilité.
  • Excellente communication et esprit d’équipe.
  • Force de proposition.

Niveau d’étude
: Formation supérieure (BAC+5) Titre d'ingénieur

Spécialité
: systèmes

Expérience souhaitée
:
4 ans d’expériences dans un poste d’administrateur Systèmes ou similaire

Permis de conduire :
déplacement à prévoir sur les différents sites de Plaine Commune
Saint-Denis , Villetaneuse et Pierrefitte-sur-Seine

Temps complet : 37h30 hebdomadaires générant 15 jours RTT annuels

Contraintes horaires:
samedi dimanche, soirées et astreintes à prévoir occasionnellement

Télétravail possible : jusqu’à 2 jours/semaine

Vos avantages :

Déplacements domicile-travail

- Prise en charge à 75% du titre de transport

- Forfait mobilité durable pour vos déplacements à vélo ou en co-voiturage

Vous bénéficiez d’un pool de véhicules et de vélos pour vos déplacements professionnels

Vous accédez à une politique sociale et de loisirs diversifiée (garde d’enfants, mutuelles, voyages, billetterie…).

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

SRE

Senior Software Engineer (AI Infrastructure)

Alignerr

Remotesenior

Apply on the employer's site

Role description

About The Role
What if your engineering skills could directly shape the platforms and systems behind the AI revolution? We're looking for Senior Software Engineers in Munich to design, build, and scale the infrastructure that powers cutting-edge AI products used by millions of people worldwide.

This is high-impact, technically challenging work — the kind that keeps great engineers engaged. You'll tackle real architectural problems, write code that matters, and collaborate with some of the sharpest minds in AI research and engineering.

This is a fully remote, flexible contract role. If you're an experienced engineer in Munich who thrives on autonomy, writes clean and scalable code, and wants to work at the frontier of technology — we want to hear from you.

  • Organization: Alignerr
  • Type: Hourly Contract
  • Location: Remote
  • Commitment: 20–40 hours/week

What You'll Do

  • Design, develop, and maintain production-grade software systems that support AI training, evaluation, and deployment pipelines
  • Write clean, well-tested, and performant code in Python, TypeScript, or other modern languages
  • Build and optimize APIs, data pipelines, and backend services that operate reliably at scale
  • Architect solutions that balance speed of delivery with long-term maintainability
  • Troubleshoot complex technical issues across distributed systems and cloud infrastructure
  • Collaborate asynchronously with cross-functional teams — including ML engineers, researchers, and product managers
  • Participate in code reviews and contribute to engineering best practices across the team
  • Work independently on well-scoped projects with meaningful ownership and autonomy

Who You Are

  • 5+ years of professional software engineering experience
  • Strong proficiency in at least one modern programming language — Python, TypeScript/JavaScript, Go, Java, or similar
  • Solid understanding of software architecture, design patterns, and system design principles
  • Experience building and deploying backend services, APIs, or data-intensive applications
  • Comfortable working with cloud platforms (AWS, GCP, or Azure) and containerized environments
  • Familiar with version control (Git), CI/CD workflows, and modern development tooling
  • Strong debugging and problem-solving skills — you can navigate ambiguity and find root causes efficiently
  • Excellent written communication skills — you can document decisions, explain tradeoffs, and collaborate effectively in an async-first environment
  • Self-directed and reliable — you manage your own time and deliver consistently without close supervision

Nice to Have

  • Experience with ML/AI infrastructure — training pipelines, model serving, data annotation platforms, or evaluation frameworks
  • Familiarity with large-scale data processing tools (Spark, Kafka, Airflow, or similar)
  • Background in distributed systems, microservices architecture, or event-driven design
  • Experience with infrastructure-as-code (Terraform, Pulumi) or Kubernetes
  • Contributions to open-source projects or a public portfolio of technical work
  • Previous experience in a remote-first or async-heavy engineering culture

Why Join Us

  • Work on cutting-edge AI infrastructure alongside world-class research labs and engineering teams
  • Fully remote and flexible — set your own schedule and work from anywhere
  • Tackle genuinely hard technical problems with real-world impact at scale
  • High autonomy and ownership — no micromanagement, just meaningful engineering work
  • Collaborate with a global community of top-tier engineers and researchers
  • Potential for ongoing work, expanded scope, and long-term contract extension as projects grow

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

SRE

IT System Engineer Microsoft 365 (m/w/d)

Rocken®

Remote

Apply on the employer's site

Role description

Digitale Verwaltung braucht solide Infrastruktur – und Profis, die sie gestalten. Unser Rocken Partner bietet professionelle ICT-Dienstleistungen für öffentlich-rechtliche Institutionen: Gemeindeverwaltungen, Schulen, Pflegeheime und weitere öffentliche Einrichtungen setzen auf seine Expertise in Sachen digitale Transformation. Das Leistungsspektrum reicht von der Konzeption bis zur Betreuung moderner Verwaltungslösungen – immer mit Blick auf Flexibilität und Zukunftsfähigkeit. Als unabhängiger Dienstleister begleitet er seine Kunden durch die Herausforderungen der Digitalisierung: von kompletten Arbeitsplatzlösungen über Hosting bis hin zu spezifischen ICT-Services. Das Team arbeitet lösungsorientiert, pragmatisch und nah am Puls der öffentlichen Hand. Hier triffst du auf echte Gestaltungsspielräume, technische Vielfalt und sinnstiftende Projekte. Bereit, öffentliche Infrastruktur digital voranzubringen? Dann bewirb dich jetzt bei unserem Rocken Partner.

Verantwortung

  • Für einen stabilen Betrieb betreust du Microsoft 365, Azure, Entra ID, Intune, Active Directory und weitere Kernsysteme
  • Du entwickelst Cloud und Infrastruktur Lösungen und begleitest deren Umsetzung
  • Die Softwarepaketierung mit PowerShell sowie deren Verteilung gehören zu deinem Alltag
  • Bei IT Projekten und der digitalen Transformation bringst du deine Expertise ein
  • Zusätzlich unterstützt du den 3rd Level Support und die Ausbildung von Lernenden

Qualifikationen

  • Fundierte Erfahrung im Microsoft-Umfeld oder eine abgeschlossene Informatikausbildung
  • Sehr gute Kenntnisse in PowerShell, M365, Azure, Active Directory und Softwarepaketierung
  • Know-how in Cloud- und ICT-Automatisierung ist von Vorteil
  • Eine selbständige, teamorientierte und lernbereite Arbeitsweise zeichnet dich aus
  • Deutsch beherrschst du schriftlich sowie mündlich einwandfrei

Benefits

  • Homeoffice
  • Attraktive Vorsorge- und Versicherungsleistungen
  • Attraktive Weiterbildungs- und Entwicklungsmöglichkeiten

ROCKEN Jobs
https://rocken.jobs

Profil Erstellen
https://rocken.jobs/application/profil\-erstellen/

  • new

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

SRE

Senior Backend Engineer, Core APIs (SRE Focus)

Jobgether

Remotesenior

Apply on the employer's site

Role description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Backend Engineer, Core APIs (SRE Focus) based in Switzerland.
This is a senior engineering role focused on building and operating highly reliable backend systems at significant scale.

You’ll develop core APIs and real-time data services where performance, availability, latency, and cost are critical design considerations.

The role combines hands-on backend engineering with strong Site Reliability Engineering ownership in production.

You’ll define and monitor service objectives, improve observability, automate operational work, and help the platform scale safely as demand grows.

You’ll take ownership of systems throughout their lifecycle, from technical design and implementation through deployment, incident response, and continuous optimization.

Working within a fully remote, globally distributed team, you’ll collaborate asynchronously while contributing to engineering standards and mentoring other developers.

This is an opportunity to solve complex distributed-systems challenges while helping build technology that protects businesses and users from online fraud.

Accountabilities

  • Design, develop, and optimize highly performant backend services for real-time data processing and web APIs, balancing reliability, latency, scalability, and infrastructure cost.
  • Define and own SLIs, SLOs, and error budgets for core API services, using reliability data to guide engineering and delivery decisions.
  • Build and maintain effective observability through meaningful metrics, structured logging, distributed tracing, dashboards, and actionable alerts.
  • Participate in an equitable on-call rotation, respond effectively to production incidents, mitigate customer impact, and lead blameless post-incident reviews.
  • Automate repetitive operational processes and follow through on incident action items to reduce operational toil and prevent recurring failures.
  • Conduct capacity planning, load testing, performance analysis, and tuning to ensure systems can scale ahead of demand.
  • Investigate complex production and product issues by developing hypotheses, running experiments, analyzing results, and translating findings into engineering improvements.
  • Integrate backend components with other services and collaborate with cross-functional teams to maintain performance, scalability, and reliability across the platform.
  • Contribute to a data-driven, reliability-focused engineering culture by sharing best practices and improving development and operational processes.
  • Mentor other engineers and support the continued growth of technical skills, engineering judgment, and operational excellence across the team.

Requirements

  • 5+ years of backend software engineering experience, including substantial experience operating production systems that you have helped build.
  • Demonstrated experience with formal SRE practices, including error budgets, capacity planning, production readiness reviews, or fault-injection and resilience testing.
  • Strong hands-on experience with Go and/or Node.js, with the flexibility and willingness to learn the other technology.
  • Proven ability to design, develop, and maintain scalable backend systems, APIs, microservices, and real-time data-processing services.
  • Practical experience running production workloads with Docker and Kubernetes or comparable containerization and orchestration technologies.
  • Experience with infrastructure-as-code tools such as Terraform or AWS CloudFormation.
  • Working knowledge of observability platforms such as Datadog, including service instrumentation, monitoring, logging, tracing, dashboards, and alerting.
  • Experience participating in production on-call rotations, troubleshooting live incidents, mitigating impact, and conducting post-incident analysis.
  • Strong SQL skills and experience working with databases or data stores such as DynamoDB, Redis, or Elasticsearch.
  • Familiarity with Git, shell scripting, IDEs, CI/CD pipelines, and standard software engineering practices.
  • Strong understanding of distributed systems, scalability, reliability, and performance engineering.
  • Excellent English communication skills and the ability to collaborate effectively within a globally distributed, remote-first environment.
  • Ability to work independently, take ownership of complex problems, and collaborate effectively across teams.
  • Strong interest in mentoring engineers and helping raise technical standards across the organization.
  • A degree in Computer Science or a related discipline is welcome but equivalent professional experience is also valued.
  • Experience with analytical platforms such as ClickHouse, Snowflake, BigQuery, Redshift, or Databricks is an advantage.
  • Experience with performance profiling and cost optimization for high-throughput cloud workloads is a plus.
  • Experience contributing to large open-source projects and working asynchronously across international teams is beneficial.
  • Experience with both Node.js and Go in backend environments, as well as knowledge of internet security and privacy mechanisms, is highly valued.

Benefits

  • US compensation reference: $150,000–$200,000 in cash compensation for US-based employees; compensation ranges vary according to hiring location.
  • Fully remote working environment with the ability to work from a wide range of countries, subject to local regulatory and security requirements.
  • Opportunity to work on high-scale backend and fraud-detection infrastructure with meaningful reliability and performance challenges.
  • Exposure to modern technologies including Go, Node.js, TypeScript, AWS, Terraform, Docker, Kubernetes, Datadog, ClickHouse, and dbt.
  • Flexible, globally distributed work environment designed around remote collaboration and asynchronous communication.
  • Opportunities for continuous technical development, experimentation, and knowledge sharing.
  • Strong emphasis on engineering ownership, reliability, mentorship, and professional growth.
  • Inclusive workplace that values diverse perspectives, experiences, and cultural backgrounds.
  • Visa sponsorship is not provided; employees must be authorized to work from their home location.
  • Country-specific employment requirements and compensation structures may apply outside the United States.

How Jobgether Works
We use an
AI-powered matching process
to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Why Apply Through Jobgether?

Data Privacy Notice:
By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

146 more openings in this category and country

Senior Azure Cloud EngineerEPAM Systems

Apply on the employer's site