Skip to content
SRE

Список недели

Все такие вакансии — одним письмом

Сейчас вы читаете одну вакансию. Таких же на витрине сотни, и каждую неделю выходят новые. Выберите, что присылать, оставьте почту — список придёт сам, без поисков и без возвращения сюда.

Считаем, сколько вышло за прошлую неделю…

Первое письмо приходит сразу, дальше — раз в неделю. Отписка в один клик из любого письма, адрес больше никуда не уходит.

SRE

Incident Manager

KX

Yerevan

Откликнуться на сайте работодателя

Описание вакансии

About KX

KX is the temporal AI infrastructure for capital markets, giving AI a sense of time, so every decision is grounded in when data was true, not just what it says. The KX portfolio normalizes market data across 250+ venues, capturing, sequencing, and replaying events at nanosecond precision. We enable firms to eliminate data fragmentation, meet regulatory auditability requirements, and build AI systems that are fast, accurate, and defensible across trading, risk, and compliance.

Trusted by 8 of the top 10 global investment banks and serving $50 trillion+ in assets under management, KX is the essential infrastructure for the world's most demanding financial institutions. We serve leading organizations across aerospace and defense and high-tech manufacturing, with offices across North America, Europe, and Asia Pacific.

Overview Of The Role

KX is hiring an Incident Manager to help streamline incident management processes, improve customer communication during critical situations, and reduce the operational pressure on technical teams.

In this role, you'll sit at the centre of customer-facing incident coordination — ensuring timely communication, effective escalation, and smooth collaboration between customers, Monitoring, Product, and Engineering teams. You'll help drive incidents from initial detection through resolution while ensuring stakeholders stay informed and aligned throughout the process.

As part of the team, you'll establish structured incident management practices, improve operational workflows, track recurring issues, and contribute to building scalable processes as we expand our customer base.

This is a high-impact role for someone who thrives in fast-paced environments, can confidently manage critical customer situations, and enjoys bringing structure and clarity to complex operational challenges.

Skills

  • Strong written and verbal English communication skills, with the ability to manage high-visibility customer communications.
  • Strong incident management and prioritization skills.
  • Ability to remain composed under pressure and handle critical customer situations.
  • Strong understanding of support SLAs and support operations principles.
  • Ability to coordinate cross-functional teams and drive issues through to resolution.
  • Deep knowledge of the Release management process.

Essential Experience

  • 3+ years of experience in incident, change, and problem management following ITIL practices.
  • Hands-on experience managing P1 incidents.
  • Experience owning customer communication during critical incidents, including updates, escalations, and closure.
  • Experience maintaining incident timelines, impact assessments, action tracking, and stakeholder communication.
  • Experience working with incident reporting, RCA documentation, and operational processes.
  • Experience working with support, monitoring, or technical teams in an operational environment.
  • Experience leading or coordinating a team of 3–5 people.

Preferred Qualifications

  • Experience in the FinTech domain
  • Awareness of Data Protection Regulations
  • Experience with Jira, Confluence, or similar collaboration tools.
  • Familiarity with monitoring and alerting tools such as Airflow, Sentry, or Grafana.

Key Responsibilities

  • Continuously monitor customer emails, channels, and threads during agreed coverage hours.
  • Together with Monitoring, identify emerging issues and notify the customers proactively. While monitoring focuses on raising the issues, incident manager focuses on customer communication side of activities in a proactive way.
  • Maintain and lead incident comms: acknowledgement → interval updates → closure → tracking remediation items and reporting on resolution status continuously.
  • Maintain incident records (timeline, impact, actions, owners, status).
  • Support monitoring team with cross- team collaboration and proactive routing of the technical work to the right team and track to closure — without owning the fix.
  • Initiate and facilitate pro-active P1 Customer calls based on needs, engaging team members required.
  • Escalate on severity, SLA risk, or customer sentiment to the PM team.
  • Track SLA breaches for incidents, escalate and speed up resolution.
  • Ensure smooth release upgrades from the incidents’ perspective, analyze the gaps, and suggest improvements to the process.
  • Managing incidents in connection with the releases as part of the release management process
  • Handle simple cases end-to-end, removing them from the PM queue.
  • In alignment with PMs and customer service tiers (priorities)improve prioritization process on the level of L1.
  • Produce incident summaries, RCA and periodic reporting+ monthly/quarterly/yearly analysis and reporting (internal and external)
  • Feed recurring patterns into monitoring coverage, run books and knowledge base.
  • Scale incident management to the whole set of Customers.
  • Participate in incident management automation via AI.
  • Drive urgent change advisory boards for incident-related patches as part of the Customer release process.

Out of scope

  • Technical resolution of issues.
  • Security and data breaches (for the beginning, later -is needed with support of the technical leads).

Location & Workplace Type

This is a Hybrid role where you will work 3 days in our Yerevan office.

Why Choose KX

Data Driven: We lead with instinct and follow facts.

Naturally Curious: We lean in, listen, and learn fast.

All In: We take ownership, take on challenges and give it our all.

Benefits

  • Competitive Salary
  • Individually tailored training and skills development
  • Private healthcare package and Employee Assistance Programme
  • Enhanced maternity and paternity package
  • Wellness Days and Volunteer Days

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

Скоро на этой странице

Резюме под эту вакансию — и билет в розыгрыш

Мы разбираем объявление до настоящих требований и переписываем ваше резюме под него — вопросами, а не выдумкой: ни одна строка не появится без вашего подтверждения. Войдите, чтобы получить это первым, — и попасть в розыгрыш.

  • Резюме под конкретную вакансию, а не «универсальное»
  • Ответы хранятся: правится любой, а не весь разговор заново
  • Всё в аккаунте — открывается с любого устройства

Разыгрываем

Скидка на сопровождение

Победителей выбираем случайно среди заявок с подтверждённой почтой. Дата розыгрыша и полные правила — на странице розыгрыша.

Правила розыгрыша

SRE

DevOps Engineer

Teza Technologies

Yerevanfull-time

Откликнуться на сайте работодателя

Описание вакансии

Объявление пришло без описания.

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

SRE

Senior Site Reliability Engineer

Wirestock

Yerevan

Откликнуться на сайте работодателя

Описание вакансии

Объявление пришло без описания.

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

SRE

AI Data Center Manager

SoftConstruct

Yerevan

Откликнуться на сайте работодателя

Описание вакансии

Are you GAME to JOIN US and be our new AI Data Center Manager?

SoftConstruct is on a search for a new player to join our AI Center Of Excellence team.

Description

We are seeking an experienced, forward-thinking AI Data Center Manager to lead our 5MW high-density compute footprint. Unlike traditional mega-scale facilities, this role is designed for an agile executive leader who can manage a cutting-edge, ultra-dense deployment focused entirely on Enterprise AI workloads, including large language model fine-tuning, high-throughput inference, and advanced multi-node HPC clustering. This role is part of a landmark AI infrastructure project with a total investment of more than USD 100 million, making it one of the largest AI data center initiatives in the region. In this role, you will hold ultimate responsibility for the physical deployment, thermal management, and day-to-day operational efficiency of our liquid-cooled GPU clusters. You will serve as the critical link between our physical colocation partners and our internal AI Platform Engineering teams to ensure maximum hardware performance, zero thermal throttling, and seamless power delivery.

Requirements

  • Professional Experience: 7+ years of direct experience in data center operations, high-performance computing (HPC), or critical facilities engineering, with at least 3 years in a direct leadership or managerial capacity
  • Advanced Thermal Expertise: Deep technical understanding of direct-to-chip liquid cooling loops, Coolant Distribution Units, rear-door heat exchangers (RDHx), and general facility MEP (Mechanical, Electrical, Plumbing) designs
  • High-Density Hardware Literacy: Direct experience deploying and maintaining high-power density compute environments (NVIDIA HGX/DGX, advanced liquid-cooled chassis, or specialized OEM architectures)
  • Colocation Contract Proficiency: Proven experience negotiating and managing SLAs, Master Service Agreements (MSAs), and joint operating procedures inside high-tier commercial colocation facilities
  • Bachelor's degree in Mechanical Engineering, Electrical Engineering, Computer Science, or equivalent technical discipline / equivalent practical experience
  • Transition Management: Experience guiding active environments through transitioning from legacy air cooling to hybrid air-and-liquid or 100% direct-to-chip liquid-cooled architectures
  • Software Infrastructure Literacy: Functional knowledge of container orchestrators (Kubernetes) and AI training scheduling architectures (Slurm)
  • Strategic Partnerships: Established, active relationships with key component providers (CDU vendors, quick-disconnect suppliers) and server OEMs

Responsibilities

  • Liquid Cooling Operations: Oversee the operation and maintenance of direct-to-chip liquid cooling loops, Coolant Distribution Units (CDUs), and localized liquid-to-air secondary heat exchangers in a dense environment
  • Rack-Level Power Strategy: Optimize active power distribution at the rack level, managing ultra-dense power envelopes ranging from 60kW to 100kW+ per rack 60-140across our 5MW footprin
  • PUE & Efficiency Optimization: Drive continuous physical optimization to lower Power Usage Effectiveness (PUE) and achieve carbon neutrality and environmental sustainability metrics
  • Colocation SLA Management: Act as primary liaison to wholesale colocation partners, holding suppliers strictly accountable to power, cooling, physical security, and infrastructure uptime SLAs
  • Hardware Lifecycle Operations: Direct the logistics, installation, staging, provisioning, and decommissioning of advanced AI hardware systems (including NVIDIA DGX/Blackwell/Hopper, Ver Rubin)
  • Network Fabric Readiness: Collaborate closely with Network Architecture teams to ensure high-bandwidth, ultra-low-latency backend topologies (InfiniBand, RoCEv2) are perfectly integrated and structurally protected
  • Telemetry and Monitoring: Implement and maintain centralized environmental telemetry pipelines that connect physical variables such as coolant flow rates, pressure, and ambient temperature with software schedulers such as Slurm and Kubernetes

Do you like to learn hard, work hard and play hard?

Do you imagine better things, technologies, future?

If you answered “Yes” to at least two of these questions then we might be a great fit for you.

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

Это все вакансии по этой категории в этой стране

Incident ManagerKX · Armenia

Откликнуться на сайте работодателя
Incident Manager — KX | mentors.coach