Skip to content
Site Reliability Engineer

Site Reliability Engineer

Staff Network Engineer (AI Fabric, Datacenter and Edge Networking) - Radian Arc (EMEA)

Submer

Remotefulltimelead

Apply on the employer's site

Role description

Location & work modality: Europe/ Remote

Start: Aug 2026

Type of Contract: Full time or Contract

About Radian Arc

Radian Arc provides an infrastructure-as-a-service (IaaS) platform for running cloud gaming, artificial intelligence and machine learning applications inside telecommunication carrier networks. Our teams across the USA, Australia, Central Europe, Malaysia, Singapore and Japan offer telecom operators a GPU-based edge computing platform without the need for capital expenditure, facilitating low latency and improved economics for value-added services and the monetization of 5G investments.

What impact you will have

Design, implement, and operate the network infrastructure powering the GPU cloud platform, including high-performance AI fabrics as well as classical datacenter networking components such as routing, security, and external connectivity. This role spans both high-performance east-west networking for distributed AI workloads and north-south connectivity, security, and inter-datacenter transport.

As the first dedicated networking role in the organization, the Staff Network Engineer combines Staff-level architectural ownership, technical direction, and cross-functional influence with hands-on execution across design, deployment, troubleshooting, automation, and operational improvement.

The Staff Network Engineer owns the long-term technical direction and operational strategy for Radian Arc's AI interconnect networks, designing scalable GPU fabrics and ensuring predictable low-latency performance across distributed training and inference workloads. The role includes designing large-scale RoCE and Ethernet fabrics, guiding architecture decisions, and ensuring operational excellence across global deployments, from hyperscale datacenters to smaller edge locations.

You will collaborate closely with platform, compute, storage, observability, and operations teams to ensure networking is deeply integrated into the overall infrastructure architecture. This role also acts as the senior escalation point for complex networking incidents, driving deep technical investigations and systemic improvements that increase reliability, latency consistency, and operational maturity across the platform.

Because this is currently the primary networking role in the company, the position is intentionally hybrid: you are expected to operate at L6 / Staff in terms of technical direction, standards, cross-team influence, and long-term design, while also directly executing critical networking work that, in a larger organization, would be distributed across multiple engineers.

What you'll do

AI Fabric & HPC Networking

  • Design and operate high-performance GPU networking fabrics supporting distributed AI workloads.

  • Architect large-scale RoCE fabrics optimized for distributed training and inference.

  • Optimize network performance for GPU communication patterns and east-west traffic.

  • Design fabric topologies such as:

  • Leaf-spine

  • Fat-Tree

  • Rail architectures

  • Multi-plane

    • Implement high-performance networking technologies including:
  • RDMA

  • RoCE

  • High-bandwidth east-west fabrics

  • Spectrum-X

  • Collaborate with compute teams to support distributed training frameworks and GPU communication libraries.

  • Define reference architectures and design principles for AI fabrics so future deployments follow reusable standards rather than one-off implementations.

  • Evaluate architectural trade-offs across performance, resilience, cost, operability, and deployment speed, and make clear recommendations to stakeholders.

Datacenter Networking

  • Design and operate Layer-2 and Layer-3 datacenter networks.

  • Implement scalable routing architectures based on BGP and ECMP.

  • Design tenant network isolation mechanisms across multi-tenant environments.

  • Implement and maintain:

  • Network bridges

  • Routing stacks

  • Overlay networking systems

  • Maintain north-south ingress/egress routing and traffic management.

  • Define standards and reusable patterns for segmentation, routing, and overlay integration across platform deployments.

Technologies include:

  • VyOS routers
  • Linux networking stacks
  • OVS / OVN
  • BGP / ECMP
  • VLAN / VRF segmentation

Security & Edge Connectivity

  • Deploy and maintain north-south security infrastructure
  • Implement WAF and application-layer protections
  • Integrate security controls with platform services

Technologies include:

  • Citrix NetScaler / Citrix WAF
  • TLS termination
  • DDoS mitigation
  • API and proxy gateway protection

Inter-Datacenter Networking

  • Design and operate private interconnects between datacenters
  • Implement and maintain dark fiber ring architectures
  • Operate high-capacity WAN connectivity between regions
  • Integrate datacenter fabrics into a global backbone network
  • Define scalable design principles for backbone evolution, inter-site routing, redundancy, and failure-domain isolation

Technologies include:

  • DWDM / dark fiber transport
  • BGP inter-site routing
  • Redundant fiber ring architectures
  • 100–400G optical transport
  • Spectrum-XGS

Engineering Execution & Delivery

  • Lead end-to-end engineering delivery of networking infrastructure, from design and labvalidation to production deployment
  • Validate network BOMs together with procurement and deployment teams
  • Provide detailed input into datacenter layouts and rack elevations
  • Drive capacity planning, performance modeling, and scaling strategies
  • Ensure network changes are executed safely with minimal customer impact
  • Act as both the architectural owner and the practical execution lead for critical network initiatives during the build-out phase of the networking function
  • Establish deployment standards, validation criteria, rollback approaches, and acceptance patterns that future engineers and teams can reuse

Operational Excellence & Reliability

  • Own operational performance and reliability of networking infrastructure

  • Drive automation for:

  • Provisioning

  • Configuration management

  • Monitoring

  • Lifecycle management

  • Improve day-2 operations through automation and operational tooling

  • Lead incident response and root-cause analysis for major network events

  • Define and track SLAs, SLOs, and reliability metrics

  • Translate major incidents and operational pain points into durable standards, design changes, and long-term architectural improvements

  • Establish measurable benchmarks for reliability, latency consistency, operability, and recovery behavior across network deployments.

Cross-Functional Collaboration

  • Work closely with infrastructure, platform, SRE, compute, storage, observability, and datacenter operations teams.
  • Provide technical leadership across infrastructure initiatives.
  • Communicate architectural decisions, trade-offs, and risks clearly to stakeholders.
  • Influence the long-term platform networking roadmap and architecture.
  • Act as the primary networking design authority across the organization, guiding adjacent teams on how networking constraints and capabilities should shape platform decisions.
  • Raise the technical bar by mentoring engineers in adjacent domains and helping build the future networking function.

Technical Stack

Datacenter Networking

  • BGP
  • EVPN / VXLAN
  • ECMP
  • VLAN / VRF
  • OVS / OVN
  • Linux networking
  • BlueField DPU

Routing & Control Plane

  • VyOS
  • BGP-based routing architectures
  • ECMP fabrics

Security

  • Citrix NetScaler / WAF
  • DDoS protection

Transport & Backbone

  • Dark fiber
  • Metro fiber rings
  • DWDM transport
  • 100–1600G optical networking

AI Networking

  • RDMA
  • RoCE
  • GPU fabrics
  • Large-scale east-west compute networking
  • Congestion control

What you'll need

Core Experience

  • Strong hands-on experience designing and operating large-scale datacenter networks

  • Expert knowledge of modern networking protocols including:

  • BGP

  • OSPF

  • ECMP

  • EVPN / VXLAN

  • Proven experience operating high-speed Ethernet networks in production environments

  • Experience operating NVIDIA / Mellanox networking platforms

  • Experience owning both architecture and direct implementation in lean or fast-scaling environments is strongly preferred

Advanced AI Fabric Networking Expertise

The candidate should have deep expertise in designing and operating networking fabrics optimized for large-scale GPU clusters and distributed AI workloads.

This includes a strong understanding of GPU communication patterns and the networking requirements of distributed training and inference systems.

Relevant expertise includes:

  • Deep understanding of NCCL communication patterns and their impact on network topology and performance.

  • Experience tuning RoCE fabrics for large-scale GPU clusters.

  • Strong knowledge of RDMA transport behavior and failure modes.

  • Practical experience implementing and tuning PFC and ECN for congestion management.

  • Understanding of GPU collective communication patterns such as all-reduce, all-gather, broadcast, reduce-scatter, and their impact on east-west network traffic.

  • Experience designing rail-optimized GPU networking fabrics for distributed training and inference clusters.

  • Familiarity with diagnosing performance issues related to:

  • NCCL stalls

  • RDMA congestion

  • Fabric hotspots

  • Packet loss impacting distributed training

  • Understanding of how networking performance affects distributed AI frameworks such as PyTorch and TensorFlow.

The candidate should also be able to collaborate closely with compute platform teams to ensure that networking infrastructure is optimized for distributed training, distributed inference, andhigh-throughput AI workloads.

Systems & Troubleshooting

  • Ability to debug complex cross-layer issues spanning:

  • Hardware

  • Firmware

  • Kernel networking

  • Distributed application communication layers

  • Strong knowledge of networking hardware, optics, and high-speed interconnects.

  • Experience designing network observability systems.

  • Strong ability to act as the senior escalation point for ambiguous, high-impact, and multi-domain technical issues.

Automation

  • Strong automation skills using Python and/or Bash.
  • Experience applying software engineering practices to infrastructure automation.
  • Experience building reusable tooling, standards, or validation approaches that increase leverage across teams.

Leadership

  • Proven ability to lead complex technical initiatives across teams.
  • Comfortable collaborating across engineering, operations, and vendors.
  • Strong systems-level thinking balancing performance, reliability, scalability, and operational cost.
  • Demonstrated ability to set architectural direction and drive adoption of engineering standards across an organization.
  • Proven ability to lead through technical influence across multiple teams and domains, without relying on formal people management authority.
  • Strong mentoring capability and ability to raise the technical level of adjacent engineering teams.
  • Able to balance short-term execution needs with long-term platform design, operational sustainability, and cost efficiency.

What we offer

  • Attractive compensation package reflecting your expertise and experience.
  • A great work environment characterised by friendliness, international diversity, flexibility, and a hybrid-friendly approach.
  • You'll be part of a fast-growing scale-up with a mission to make a positive impact, offering an exciting career evolution.

Our job titles may span more than one job level. The actual base pay is dependent on a number of factors, such as transferable skills, work experience, business needs and market demands.

Our inclusive responsibility

Radian Arc is committed to creating a diverse and inclusive environment and is proud to be an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, veteran status, or any other protected category under applicable law.

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Coming to this page

A resume for this role — and a ticket to the draw

We take the posting apart down to the real requirements and rewrite your resume against it — by asking, not inventing: no line appears without your confirmation. Sign in to get it first, and to enter the draw.

  • A resume for this exact role, not a universal one
  • Answers are kept: edit any one, not the whole conversation
  • All in your account — open it from any device

On the wheel

A discount on mentoring

Winners are drawn at random among entries with a confirmed email. The date and the full rules are on the draw page.

Draw rules

Site Reliability Engineer

Ingénieur Système, Sauvegarde et réseaux à Paris H/F

Free-Work

Remotefulltimesenior

Apply on the employer's site

Role description

Dans le cadre du développement de notre équipe IT chez l'un de nos clients grands comptes, nous recherchons un(e) Ingénieur(e) Système Sauvegarde réseaux H/F afin d'assurer l'administration, la maintenance et l'évolution des infrastructures systèmes de nos environnements clients.

Vos missions seront les suivantes :

  • Participer aux projets d'évolutions de la plateforme technique de la Video Factory ( tête de réseau OTT et IPTV
  • Conception / participation au POC avec le N3 (les experts) / intégration / ingénierie / recette unitaire / recette des workflow et documentation des briques techniques et des procédures.
  • Assurer la maintenance en condition opérationnelle de la plateforme
  • Organisation des sauvegardes et des montées de version des équipements IT (VM, NAS, OS (linux et windows)...) et Vidéos (encodeurs, DCM, serveurs d'origine, sondes ) de la plateforme.
  • Analyse des risques et des impacts potentiels et planification en HNO le cas échéant.
  • Assurer le « run » de la plateforme en tant que support niveau 2 en soutien des équipes support de niveau 0 et 1
  • Analyse d'incidents et résolutions, escalade au N3 et aux fournisseurs (ouverture et suivi des tickets) le cas échéant, communication sur les avancées les plus significatives et les impacts majeurs,
  • Organiser les activités des fournisseurs (mises à jour et évolution technique),
  • Assurer le suivi des déploiements et mettre en place les contrôles (recette).
  • Participer à l'amélioration de l'organisation du support global de la plate-forme technique :
  • Formation des équipes de maintenance, documentation des procédures d'exploitation.

     Des missions complémentaires peuvent être confiées.

Référence de l'offre : cmk6d3cswx

Profil candidat:

Profil recherché :

Issu(e) d'une formation supérieure en informatique (BAC+5 / Diplôme d'ingénieur),

Vous justifiez d'au moins 5 ans d'expérience sur un poste similaire.

Compétences techniques souhaitées :

  • Maîtrise des environnements Windows Server, Linux et Vmware (des connaissances sur Nutanix est un plus).
  • Bonne connaissance des environnements vidéo et audio sur IP.
  • Connaissance des réseaux IP et de l'adressage multicast.
  • Maitrise des outils d'analyse de qualité vidéo ainsi que des outils de supervision.
  • Bonnes capacités d'analyse et de résolution de problèmes.
  • Autonomie, rigueur et bon relationnel.
  • Capacité à travailler en équipe et à intervenir dans des environnements de production

Environnement technique :

DCM, Anevia, Imagine, Nevion, Harmonic xOS, OpenHeadEnd, Elemental, USP.

Produits systèmes, virtualisation : VMware, Wallix, Nutanix, NAS, Debian, Windows Server, FTP

Ce que nous vous proposons :

Valeurs : en plus de nos 3 fondamentaux que sont l'audace, la bonne foi et la réactivité, nous garantissons un management à l'écoute et de proximité, ainsi qu'une ambiance familiale.

Contrat : CDI ou Freelance

Localisation : Paris

Package rémunération & avantages :

  • Le salaire : rémunération annuelle brute selon profil et compétences
  • Les basiques : mutuelle familiale et prévoyance, titres restaurant, remboursement transport en commun à 50%, avantages du CSE (culture, voyage, chèque vacances, et cadeaux), RTT (jusqu'à 12 par an), plan d'épargne, prime de participation
  • Nos plus : forfait mobilité douce & Green (vélo/trottinette et covoiturage), prime de cooptation de 1000 € brut, e-shop de matériel informatique à des prix préférentiels (smartphone, tablette, etc.)
  • Votre carrière : plan de carrière, dispositifs de formation techniques & fonctionnels, passage de certifications, accès illimité à Microsoft Learn
  • La qualité de vie au travail : télétravail avec indemnité, évènements festifs et collaboratifs, accompagnement handicap et santé au travail, engagements RSE

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

SAP DevOps Engineer (f/m/d)

E.ON Digital Technology

Remote

Apply on the employer's site

Role description

You have a passion for technology and want to make the world a greener place?
Then become a playmaker (f/m/d) and join our team as SAP DevOps Engineer (f/m/d) at E.ON Digital Technology.

We play a key role in shaping the energy transition by leading E.ON's digital transformation across Europe. We explore new paths by developing ideas, breaking new ground, making visions reality, and bringing new technologies to life. We deliver sustainable technology solutions because…

… it’s on us to make new energy work!
The Team
– your impact

At E.ON, the SAP Engineering Chapter is a collective of innovative minds dedicated to delivering world-class SAP architecture, SAP software development and SAP engineering capabilities across our segments and product teams. By joining us, you will play a critical role in keeping E.ON's SAP architecture and capabilities modern, secure and excellent.

Your Role –
meaningful & rewarding

As a SAP DevOps Engineer (f/m/d) at E.ON, you will be responsible for designing, developing, training and maintaining SAP DevOps solutions for our SAP system landscape and platforms. You will work closely with the SAP DevOps teams and developers, and other stakeholders to ensure the provisioning of a modern, user-friendly, state-of-the-art SAP DevOps pipeline.

  • Design, implement and continuously improve DevOps concepts, CI/CD templates, pipelines and automation for SAP landscapes (e.g. SAP BTP, S/4HANA)
  • Enable reliable build, test and deployment processes across SAP Cloud and on-premise environments
  • Establish observability concepts for SAP BTP and on-premise applications for monitoring, health checks, alerts and APM by using tools such as SAP Cloud ALM, New Relic and Uptrends
  • Collaborate closely with SAP development, architecture, security and operations teams
  • Conduct DevOps maturity assessments and guide teams on their DevOps Journey
  • Ensure high availability, performance, security and compliance of SAP systems
  • Troubleshoot complex problems and support root cause analysis
  • Continuously evaluate new SAP DevOps tools, technologies and best practices

Your Profile
– authentic & open-minded

  • Strong experience in DevOps, platform engineering or operations within SAP environments
  • Hands-on experience with CI/CD and security, as well as quality tooling (e.g. GitLab CI/CD, SonarQube, Renovate or similar)
  • Strong knowledge in development, automated testing and testing strategies
  • Experience with scripting and automation (e.g. Bash, Groovy)
  • Solid knowledge of SAP S/4HANA, SAP BTP and related deployment and transport mechanisms
  • Knowledge of Kubernetes and container technologies and IaaC principles is a plus
  • Strong understanding of security, monitoring and reliability concepts
  • Structured, proactive and solution-oriented working style
  • Fluent in English

Our Benefits
– smart & useful

  • Advance your development: We grow and we want you to grow with us. Learning on the job, exchanging with others, or taking part in an individualtraining – our learning culture enables you to bring your personal and professional development to the next level.
  • Recharge your battery: You have 30 days of paid vacation per year plus Christmas and New Year's Eve off. Your battery still needs charging? You canexchange parts of your salary for more paid vacation or you can take a sabbatical.
  • Enjoy hybrid work: We combine office collaboration with focused work from home. It’s also possible to go on workation for up to 20 days per year withinEurope.
  • Stay active & healthy: Benefit from a company-sponsored health membership.
  • Elevate your mobility: From car and bike leasing offers to a subsidised Deutschland-Ticket – your way is our way.
  • Think ahead: With our company pension scheme and a great insurance package we take care of your future.
  • This is by far not all… We are looking forward to speaking with you about further benefits during the hiring process.

Do you have questions?
For further information please contact Agneta Lierl, EDT_Talent_Acquisition@eon.com.

What you need to know:
Contract type: Permanent

Working time: Full time

Company: E.ON Digital Technology GmbH

Location: Essen, Hannover, München, Berlin, Würzburg, Hamburg, Frankfurt am Main

Function area: IT/Digital; Engineering

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

Senior Security Engineer (m/w/d)

Rocken®

Remotesenior

Apply on the employer's site

Role description

Steuerverwaltungen brauchen clevere Software – und dahinter stehen clevere Köpfe. Unser Rocken Partner entwickelt die führende Business Lösung für kantonale und kommunale Steuerverwaltungen. Die Anwendung deckt den kompletten Verwaltungsprozess ab: vom Steuerregister über Veranlagungen und Fakturierung bis zum Inkasso und zur Verlustscheinbewirtschaftung. Derzeit entsteht eine neue Software-Generation – ein spannendes Projekt, das technisches Know-how und Innovationskraft vereint. Die Arbeitsweise ist geprägt von überlegter Fokussierung, lebendigem Austausch zwischen Teams und intellektueller Courage. Elegante Lösungen entstehen durch Wissen, Geist und Teamenergie. Bereit, an einem Produkt zu arbeiten, das echten Impact hat? Unser Rocken Partner sucht Menschen, die sich für eine gemeinsame Idee begeistern und die digitale Zukunft der Steuerverwaltung mitgestalten möchten.

Verantwortung

  • Du analysierst Schwachstellen und setzt passende Sicherheitslösungen um.
  • Du arbeitest an Themen wie Zugriff, Netzwerk, Pentesting und Monitoring.
  • Du entwickelst die Sicherheitsarchitektur weiter und beachtest ISO 27001.
  • Du automatisierst in Windows/Linux und unterstützt Kubernetes-Setups.

Qualifikationen

  • Du hast eine IT-Ausbildung und Erfahrung als System Engineer mit Security-Fokus.
  • Du denkst vernetzt, erkennst Risiken und arbeitest agil mit modernen Tools.
  • Du kennst Azure, VMware und gängige Security-Lösungen.
  • Du brennst für IT-Security und entwickelst im Team nachhaltige Lösungen.

Benefits

  • Interessante und abwechslungsreiche Tätigkeiten/Projekte
  • Attraktive Weiterbildungs- und Entwicklungsmöglichkeiten
  • Flexible Arbeitszeitgestaltung
  • Homeoffice
  • Offene Unternehmenskultur
  • Beteiligung oder Übernahme ÖV-Abonnements
  • Beteiligung oder Übernahme Parkplatz
  • Attraktive Mitarbeiterrabatte
  • Kostenlose Früchte und Getränke

ROCKEN Jobs
https://rocken.jobs

Profil Erstellen
https://rocken.jobs/application/profil\-erstellen/

  • new

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

Microsoft Cloud & System Engineer (m/w/d)

Rocken®

Remote

Apply on the employer's site

Role description

Unser Rocken® Partner ist spezialisiert auf IT-Security, Cloud Lösungen, sowie IT-Outsourcing und Support. Sie setzen auf innovative Services und Dienstleistungen, bewegen und orientieren sich am Puls der Technik. Daraus ergeben sich für unsere Kunden viele Vorteile gegenüber den traditionellen IT-Lösungen, die weniger flexibel und selten skalierbar sind.

Verantwortung

  • Sicherstellung einer stabilen IT-Infrastruktur und hohen Systemverfügbarkeit
  • Betreuung und Weiterentwicklung von Microsoft-Umgebungen im Kundenumfeld
  • Planung und Umsetzung von Infrastruktur-, Rollout- und Migrationsprojekten
  • Analyse und Behebung technischer Störungen im 1st- und 2nd-Level-Support
  • Direkter Kundensupport, technische Beratung und Betreuung vor Ort

Qualifikationen

  • Microsoft 365, Azure, Entra ID und Microsoft Intune
  • Kenntnisse in Client-/Server-Systemen, Virtualisierung, Monitoring und Backup
  • Erfahrung im 1st-/2nd-Level-Support und technischen Troubleshooting
  • Kenntnisse in Infrastrukturplanung, System-Rollouts und Migrationen
  • Sehr gute Deutschkenntnisse sowie gute Englischkenntnisse; weitere Sprachen von Vorteil

Benefits

  • Flexible Arbeitszeitgestaltung
  • Homeoffice
  • Zahlreiche Mitarbeiterevents
  • Beteiligung oder Übernahme Parkplatz
  • Kostenlose Früchte und Getränke
  • Attraktive Weiterbildungs- und Entwicklungsmöglichkeiten

ROCKEN Jobs
https://rocken.jobs

Profil Erstellen
https://rocken.jobs/application/profil\-erstellen/

  • new

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

236 more openings in this category and country

Staff Network Engineer (AI Fabric, Datacenter and Edge Networking) - Radian Arc (EMEA)Submer

Apply on the employer's site