Skip to content
Site Reliability Engineer

The week's list

Every role like this one, in one letter

You are reading one posting. There are hundreds like it on the board, and new ones every week. Pick what you want, leave an email, and the list comes to you — no searching, no coming back here.

Counting what came out this past week…

The first letter arrives right away, then one a week. Unsubscribe in one click from any letter — the address goes nowhere else.

Site Reliability Engineer

Sr. Cloud Operations Engineer

Tungsten Automation

Sofiafulltimesenior

Apply on the employer's site

Role description

Job Purpose
The Senior Cloud Operations Engineer, Cloud Services supports and manages of the software components of the Tungsten Automation Cloud Solutions technology stack to deliver services to customers. Engineers are expected to provide first-class system operation and support to the Tungsten Automation Software cloud enterprise. The senior engineer must demonstrate technical maturity beyond that of the engineer and be able to professionally communicate with team members, management, other department personnel, third-party vendors, customers, and partners. They must work effectively in a team environment and provide guidance to and share expertise with others in the team. They must work effectively in a team environment. Successful candidates must demonstrate the ability to apply technology to support the needs of Cloud Solutions customer and partners.

Key Responsibilities
Specific responsibilities include, but are not limited to, the following:

Application Operation Duties

  • Respond to requests from other support teams for incident support to manage and resolve incidents requiring intervention at platform process and application levels.
  • Technical Application and Customer deployment of software and infrastructure

Cloud Platform Operations

  • Monitor system status, activity, performance metrics and capacity
  • Participate in weekly on call rotation to support the incident management process
  • Ensure system patches and updates are kept current
  • Be first line of contact for daily operations towards our vendors
  • Maintain up-to-date records of all system assets and licenses
  • Manage cloud technology resources such as compute and storage
  • Monitor the elasticity of cloud solutions and report irregularities and identify possible incidents
  • Monitor vendor performance periodically to ensure that service levels and commitments are being met by both sides

Team Specific Duties

  • Act as first among equals and subject matter expert for management of cloud operations
  • Deputize as required for the Cloud Operations Manager
  • Support and, where appropriate, mentor less experienced members of the team
  • Utilize AI-enabled tools (e.g., chatbots, documentation automation, analytics assistants) to improve efficiency, accuracy, and streamline routine tasks while following company AI governance and data privacy standards

Base salary range: For this role, based in Bulgaria, the base salary range is €32500-52000 per year, set using objective, gender-neutral job evaluation criteria applied to all employees performing this role or work of equal value.

Determining your actual pay: Your base salary within this range will be based on objective, gender-neutral factors — work location, skills, qualifications, experience, and education/training. It will not be reduced based on prior salary or negotiating position, nor will it fall below the range or below what colleagues in equal work receive.

Scope: This range covers base salary only, not bonuses, benefits, or other remuneration (as applicable).

Your rights: Tungsten Automation does not ask about current or past pay, and you need not disclose it. This notice does not limit your right to negotiate or your entitlement to equal pay for equal work or work of equal value.

While the job description describes what is anticipated as the requirements of the position, the job requirements are subject to change based upon any changing needs and requirements of the business.
Required Skills

  • Monitor system status, activity, performance metrics and capacity
  • Participate in weekly on call rotation to support the incident management process
  • Ensure system patches and updates are kept current
  • Be first line of contact for daily operations towards our vendors
  • Maintain up-to-date records of all system assets and licenses
  • Manage cloud technology resources such as compute and storage
  • Monitor the elasticity of cloud solutions and report irregularities and identify possible incidents
  • Monitor vendor performance periodically to ensure that service levels and commitments are being met by both sides
  • Skills in prompting AI systems and assessing output quality
  • Ability to leverage AI to ideate, develop and scale to the needs of the department

Required Experience

  • Minimum of 4 years’ proven expertise on both AWS and Azure public cloud services
  • Minimum of 5 years’ experience with Linux and or Windows Server
  • Microsoft Azure and Amazon Web Services technology
  • Experience of working in a Cloud/Software as a Service environment
  • Experience with SQL Server and other database technologies
  • Comprehensive understanding of networking protocols, technologies, and LAN/VLAN configurations
  • Strong technical, analytical, and problem-solving skills
  • Excellent written, verbal, and organizational skills
  • Ability to demonstrate skills relevant to incident and problem management
  • Experience as the principal SME within other technical teams
  • Escalate support with outside consultants/vendors when necessary
  • Proven experience of having transferred knowledge to and supported other team members
  • Prioritize and manage multiple tasks while remaining detail-oriented
  • Communicate, in writing and orally, in a professional fashion
  • Accurately and professionally document all work

Tungsten Automation Corporation is an Equal Opportunity Employer M/F/Disability/Vets

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Coming to this page

A resume for this role — and a ticket to the draw

We take the posting apart down to the real requirements and rewrite your resume against it — by asking, not inventing: no line appears without your confirmation. Sign in to get it first, and to enter the draw.

  • A resume for this exact role, not a universal one
  • Answers are kept: edit any one, not the whole conversation
  • All in your account — open it from any device

On the wheel

A discount on mentoring

Winners are drawn at random among entries with a confirmed email. The date and the full rules are on the draw page.

Draw rules

Site Reliability Engineer

Linux System Administrator

FLOLIVE®

Sofia

Apply on the employer's site

Role description

floLIVE is rewriting the playbook for the global IoT Connectivity landscape. Our groundbreaking Connectivity Management Service is reshaping the way Enterprises, Cloud providers, IoT service providers, and Mobile Operators connect and manage their devices across the globe.

Our global presence includes the United Kingdom, United States, Israel, Bulgaria, and China, solidifying our commitment to local, secure, and compliant connectivity solutions. With our innovative software-defined connectivity technology, FloLIVE delivers seamless connectivity management across boundaries.

Become part of the IoT revolution as we redefine what's possible.

Position Overview:

floLIVE is looking for a Linux System Administrator, who will be in charge of architecting, installing, configuring, and optimizing computing systems and networks in a Linux-based environment.

Responsibilities:

  • Install, configure, and maintain hardware and software components in a virtualized Linux Environment.
  • Fully manage the deployment, configuration, and troubleshooting of all aspects of Linux servers.
  • Assess, approve, and administer all equipment, hardware, and software upgrades.
  • Provide direct mission support in a classified and unclassified Linux desktop and server environment.
  • Administer, troubleshoot, and resolve failures and critical alerts on Linux and Windows servers.
  • Write and troubleshoot Windows and Linux Shell Scripts
  • Monitor and maintain NAS storage.
  • Monitor the performance of our computer systems, and proactively identify potential issues.

Qualifications

:

  • Excellent knowledge of Linux (Ubuntu, CentOS, Redhat);
  • Familiarity with virtualization platforms such as VMWare, Proxmox, or any KVM-based deployments;
  • Good understanding of current protocols and standards, including Linux security, Cloud security, Core Switching/Routing, SSL/IPSec, SAN, and Virtualization;
  • Hands-on experience troubleshooting hardware such as servers, routers, switches, network interface cards, and so on.
  • Knowledge and understanding of system flow charts, data processing concepts, and telecommunications principles.
  • Good understanding of best practices in Data Center operations and support
  • Experience with backup and disaster recovery principles.
  • Background in SaaS infrastructure and service development, ideally across a wide range of software, operational, and networking aspects.
  • Work independently functioning in a team environment.
  • Knowledge of technologies like Pacemaker/Kubernetes
  • MySQL - advantage

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

BMC Helix/Remedy AIOps Engineer - Automation & Autonomous Operations (m/f/d)

CGI

Sofiafulltime

Apply on the employer's site

Role description

We’re establishing a brand-new technology and innovation hub in Sofia – and you can be part of it from day one.

Here, you won’t just be shaping our services for German clients; you’ll also help us build a strong presence in the Bulgarian market. You’ll collaborate closely with our teams in Germany, bring in your ideas, and influence the tools, processes, and culture of our new Sofia location from the ground up.

In short:
We’re blending the energy of a fresh start with the reliability of a global player. If you’re excited about building something new both locally and internationally, this is the place to make your mark.
Apply now and be part of our success story in Sofia! #BePartofourStory Let‘s #GrowTogetherNow

Your future duties and responsibilities:

You will join a team focused on modernizing IT operations and moving toward more automated, AI-driven ways of working. You will work with BMC Helix (formerly BMC Remedy), combining ITSM, monitoring, event management, observability, and automation to build more intelligent and increasingly autonomous IT operations. Your role will combine hands-on engineering with platform design and continuous improvement, with the long-term goal of enabling closed-loop and increasingly autonomous IT operations.

  • Design, implement, and evolve BMC Helix and AIOps solutions, covering event management, correlation, anomaly detection, root-cause analysis, service topology, and operational automation
  • Integrate ITSM processes with monitoring and automation, connecting Incident, Problem, and Change Management with observability platforms and operational workflows
  • Improve operational efficiency and reduce MTTA/MTTR through event enrichment, filtering, correlation, service mapping, topology modeling, and automated remediation
  • Build and integrate automation workflows using tools such as Ansible, Rundeck, StackStorm, Jenkins, APIs, webhooks, and integration/iPaaS platforms
  • Support observability and reliability practices across cloud and on-premises environments, including monitoring, health dashboards, SLIs/SLOs, service dependencies, and operational insights
  • Help introduce AI-first capabilities into IT operations, including incident summarization, knowledge-based assistance, intelligent recommendations, and agentic automation, with appropriate governance and safeguards

Required Qualifications To Be Successful In This Role:

  • 4+ years of experience in IT Operations, SRE, DevOps, Platform Engineering, Monitoring/Observability, ITSM, or a similar technical role
  • Hands-on experience with BMC Helix (formerly BMC Remedy), BMC Helix Operations Management/AIOps, or comparable enterprise ITSM and operations management platforms
  • Good understanding of ITIL/ITSM processes, particularly Incident, Problem, and Change Management, and experience integrating ITSM with monitoring or operational tooling
  • Experience with monitoring, event management, observability, CMDB/discovery, or service mapping technologies, such as BMC solutions, Prometheus, Grafana, Elastic, Splunk, Dynatrace, Datadog, New Relic, or similar
  • Practical automation and integration skills using Python and/or Bash, REST APIs, JSON/YAML, Git, and automation/orchestration tools such as Ansible, Rundeck, Jenkins, or equivalent
  • Experience working with Linux and enterprise infrastructure, ideally including containers, Kubernetes/OpenShift, cloud or hybrid environments, networking, TLS/certificates, and integration platforms. Excellent English communication skills are required; German is considered an advantage

What We Offer:

  • Team Culture: You’ll find colleagues here who make collaboration enjoyable. We interact openly, use first names across all levels, and don’t think in hierarchies or silos
  • Contract & Working Time: Permanent, full-time contract with a standard 8-hour workday
  • Flexible Work Setup: Enjoy the best of both worlds with our hybrid work model – combining office presence in Sofia with remote work from home
  • Working Hours: Flexible working hours tailored to project needs and client requirements
  • Leave Policy: 25 days of paid vacation + 1 day for your Birthday
  • Premium supplementary health insurance package: Including coverage for prescription glasses, reimbursement of medical expenses, and dental care
  • Volunteer Opportunities: Make a positive impact with 8 hours of paid leave for voluntary work
  • Career Development: Benefit from clear career progression paths and opportunities to grow within the company
  • Continuous Learning: Access our internal learning platform, CGI Academia, to shape your skills based on your interests. Learn up to 9 languages via goFLUENT
  • Referral Rewards: Extend your network – refer friends and earn bonuses through our rewarded referral program
  • Health & Wellbeing: Access our Health and Wellbeing Portal for resources that support your physical and mental health
  • Multisport Benefit: Аvailable at a preferential price for access to various sports and wellness facilities
  • Office Environment: Work in a state-of-the-art co-working space in Sofia, with complimentary coffee, tea, and water to keep you refreshed throughout the day

Together, as owners, let’s turn meaningful insights into action.

Life at CGI is rooted in ownership, teamwork, respect and belonging. Here, you’ll reach your full potential because…

You are invited to be an owner from day 1 as we work together to bring our Dream to life. That’s why we call ourselves CGI Partners rather than employees. We benefit from our collective success and actively shape our company’s strategy and direction.

Your work creates value. You’ll develop innovative solutions and build relationships with teammates and clients while accessing global capabilities to scale your ideas, embrace new opportunities, and benefit from expansive industry and technology expertise.

You’ll shape your career by joining a company built to grow and last. You’ll be supported by leaders who care about your health and well-being and provide you with opportunities to deepen your skills and broaden your horizons.

That same commitment to fairness extends to how we use technology. To support our recruitment team, AI tools may be used to help assess applications though they never replace human judgement. All hiring decisions remain entirely in the hands of our recruitment professionals.

Come join our team—one of the largest IT and business consulting services firms in the world.

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

Senior Site Reliability Engineer (SRE)

Mirantis

Sofiasenior

Apply on the employer's site

Role description

About Mirantis
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment—on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy.

Job Description
We are looking for a highly experienced and driven Senior Site Reliability Engineer to join our forward-thinking cloud development and operations team. In this role, you will contribute to the design, development, and operation of sophisticated cloud-based AI solutions built on the CNCF ecosystem including Kubernetes, running on cutting-edge hardware from leading vendors. This role focuses mainly on deploying AI infrastructure built on NVIDIA-certified hardware, following architecture and implementation designs produced by our engineering team. You will play a pivotal role in ensuring the reliability, security, and performance of container infrastructure, while mentoring team members and Mirantis customers to deliver high-quality software and services. As a senior engineer, you will work closely with stakeholders to define technical strategies, solve complex challenges, and ensure the seamless integration of cloud and software services. This is an excellent opportunity to make a significant impact while driving innovation in a rapidly evolving cloud ecosystem.

Main Responsibilities:

  • Work with geographically distributed international teams on technical challenges and process improvements.
  • Develop, implement, maintain, and troubleshoot cloud and AI infrastructure solutions based on open source software.
  • Collaborate with stakeholders to gather and refine technical requirements.
  • Optimize system performance, reliability, and scalability.
  • Troubleshoot, debug, and resolve complex technical issues.
  • Participate in code reviews to maintain high quality standards.
  • Stay up to date with industry trends and best practices in cloud operations and development.
  • Design and implement AI-driven automation across the DevOps lifecycle, including code development and maintenance.
  • Facilitate knowledge transfer to customers during the delivery phases.

Qualifications

  • 5+ years of professional experience in DevOps, with a strong focus on cloud and infrastructure technologies, including Kubernetes and/or OpenStack
  • Experience with high-performance data center processing, networking, and storage
  • Exposure to Golang and working knowledge of other programming languages (Python, JavaScript).
  • Strong knowledge of distributed systems, microservices architecture, and CI/CD pipelines.
  • Exceptional problem-solving and debugging skills across networking and storage (hardware and software), Linux, and Kubernetes, with attention to performance optimization and security.
  • Demonstrated ability to lead technical tasks and collaborate effectively with diverse teams.
  • Comfortable making independent judgment calls when working directly with customers, often with limited day-to-day oversight.
  • Excellent written and spoken English.
  • Excellent customer-facing communication skills.
  • A commitment to innovation, continuous learning, and delivering high-quality results.
  • Ability to travel up to 25% if needed, including internationally.

Nice to have

  • Extensive experience in network and/or storage architecture.
  • Experience with high-performance computing or GPU infrastructure: GPU scheduling, MIG/vGPU, RDMA/RoCE or InfiniBand fabrics, NVLink, DCGM health-checking, GPU driver/firmware lifecycle or NVIDIA AI Enterprise.
  • Presence in the open source community including upstream contribution and conference presentations.
  • Prior experience with commercial container and virtual compute infrastructure platforms such as Rancher, Openshift, and VMware.

Education and Experience:

  • Bachelor's degree in Computer Science or a related field, or equivalent experience.
  • At least 5 years of DevOps or Software Development experience or in a similar role.

Additional Information
What does Mirantis offer you?

  • Work with an established Silicon Valley leader in the cloud infrastructure industry;
  • Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;
  • Be a part of cutting-edge, open-source innovation;
  • Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued;
  • Professional development and training;
  • Attend conferences and working groups;
  • Company outings, happy hours, hackathons, and tech talks;
  • Receive a competitive compensation package with a strong benefits plan.

It is understood that Mirantis, Inc. may use automated decision-making technology (ADMT) for specific employment-related decisions. Opting out of ADMT use is requested for decisions about evaluation and review connected with the specific employment decision for the position applied for. You also have the right to appeal any decisions made by ADMT by sending your request to isamoylova@mirantis.com
By submitting your resume, you consent to the processing and storage of your personal data in accordance with applicable data protection laws, for the purposes of considering your application for current and future job opportunities.

We are a Leader for Container Management in G2 (#2 after AWS)!

We are a Leader for Container Management in G2 (#2 after AWS)!

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Site Reliability Engineer

Site Reliability Engineer

Man Group

Sofiafulltime

Apply on the employer's site

Role description

About Man Group
Man Group is a global alternative investment management firm focused on pursuing outperformance for sophisticated clients via our Systematic, Discretionary and Solutions offerings. Powered by talent and advanced technology, our single and multi-manager investment strategies are underpinned by deep research and span public and private markets, across all major asset classes, with a significant focus on alternatives. Man Group takes a partnership approach to working with clients, establishing deep connections and creating tailored solutions to meet their investment goals and those of the millions of retirees and savers they represent.

Headquartered in London, we manage $227.6 billion* and operate across multiple offices globally. Man Group plc is listed on the London Stock Exchange under the ticker EMG.LN and is a constituent of the FTSE 250 Index. Further information can be found at www.man.com

  • As at 31 December 2025

The Role
Join our high-performing Site Reliability Engineering (SRE) team and play a pivotal role in ensuring the reliability, scalability, and performance of the technology powering Man Group’s hedge funds. You’ll have the autonomy, tools, and support to innovate and shape the future of our platform. This is an opportunity to work on cutting-edge projects, gain mentorship from senior leaders, and develop a deep understanding of both technology and the business.

As an SRE, you’ll take ownership of service reliability and deliver solutions that make a real impact. Your initial focus will include leveraging AI to accelerate incident diagnosis and resolution, improving observability, capacity planning, and automation. Over time, you’ll work across our entire infrastructure stack, operating at scale and driving continuous improvement.

Role Responsibilities

  • Ensure reliability and performance of critical systems across global infrastructure through proactive monitoring and rapid incident response.
  • Design and implement observability solutions using tools like Prometheus, Grafana, ELK, and Loki to provide deep insights into system health.
  • Automate operational tasks and build self-service capabilities to eliminate toil and improve efficiency.
  • Develop and maintain SLIs, SLOs, and error budgets to guide reliability improvements and inform engineering priorities.
  • Participate in incident response efforts, blameless post-mortems, and implement preventive measures to reduce recurrence.
  • Collaborate with development teams to improve system design, deployment practices, and operational excellence.
  • Operate at scale, managing petabyte-level storage, large CPU/GPU deployments, and high-throughput distributed systems.
  • Contribute to capacity planning and performance tuning, ensuring systems meet business demands.
  • Manage multiple ELK clusters hosting hundreds of terabytes of logs, telemetry, and APM data.

Key competencies
Required

  • Strong understanding of SRE principles, including SLIs, SLOs, error budgets, and reliability best practices.
  • Hands-on experience with observability and monitoring tools (Prometheus, Grafana, ELK, Loki, or similar).
  • Proficiency with automation tools (Ansible, Terraform) and scripting/programming languages (Python, Go, PowerShell).
  • Strong troubleshooting and debugging skills across distributed systems, with the ability to diagnose complex production issues under pressure.
  • Experience with incident management, on-call rotations, and post-incident reviews.
  • Familiarity with Kubernetes and container orchestration.
  • A proactive mindset and ability to take ownership of reliability initiatives.

Advantageous

  • Experience with CI/CD pipelines and source control workflows (Git, Jenkins, TeamCity).
  • Administration of Linux and Windows systems and exposure to cloud technologies (AWS/Azure).
  • Understanding of networking concepts, load balancing, and distributed architectures.
  • Knowledge of AI/LLM concepts (context windows, prompt tuning, MCP servers).
  • Interest in FinOps principles, desire to understand the true cost of our decisions.
  • Excellent communication and collaboration skills.

Benefits

  • Modern office located in the OfficeX campus with easy access to transport and amenities.
  • Hybrid working model
  • Competitive compensation package
  • 25 days holiday allowance
  • Premium Health insurance
  • Employee Assistance program
  • Referral Bonus
  • Additional days off for long service and volunteering
  • Multisport card
  • Opportunities for professional development including internal tech talks
  • Conference attendance, and engagement with the open-source community.

Inclusion, Work-Life Balance And Benefits At Man Group
You'll thrive in our working environment that champions equality of opportunity. Your unique perspective will contribute to our success, joining a workplace where inclusion is fundamental and deeply embedded in our culture and values. Through our external and internal initiatives, partnerships and programmes, you'll find opportunities to grow, develop your talents, and help foster an inclusive environment for all across our firm and industry. Learn more at www.man.com/diversity.

You'll have opportunities to make a difference through our charitable and global initiatives, while advancing your career through professional development, and with flexible working arrangements available too. Like all our people, you'll receive two annual 'Mankind' days of paid leave for community volunteering.

Our comprehensive benefits package includes competitive holiday entitlements, pension/401k, life and long-term disability coverage, group sick pay, enhanced parental leave and long-service leave. Depending on your location, you may also enjoy additional benefits such as private medical coverage, discounted gym membership options and pet insurance.

Equal Employment Opportunity Policy
Man Group provides equal employment opportunities to all applicants and all employees without regard to race, color, creed, national origin, ancestry, religion, disability, sex, gender identity and expression, marital status, sexual orientation, military or veteran status, age or any other legally protected category or status in accordance with applicable federal, state and local laws.

Man Group is a Disability Confident Committed employer; if you require help or information on reasonable adjustments as you apply for roles with us, please contact TalentAcquisition@man.com.

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

22 more openings in this category and country

Sr. Cloud Operations EngineerTungsten Automation · Bulgaria

Apply on the employer's site