Skip to content
SRE

SRE

Monitoring & Observability Engineer

SkyHighGrowth Inc.

Belgrade

Откликнуться на сайте работодателя

Описание вакансии

We are looking for an Monitoring & Observability Engineer to support the design and build-out of a modern observability and event management capability, with a strong focus on Datadog, ServiceNow integration, monitoring policy migration, and event pipeline modernization.

This role will be responsible for extracting, analyzing, and re-mapping existing BMC TrueSight/Helix monitoring policies and event rules into a new observability model. The ideal candidate has hands-on experience building or significantly improving monitoring and event management environments, not just maintaining existing tools.

This person will play a
key role in designing log and metric pipelines, supporting SNMP trap processing, implementing Datadog-ServiceNow integration
, and helping define a tiered observability strategy that balances production-critical visibility with cost-efficient lower-tier monitoring using OpenTelemetry-based open-source pipelines.

Key responsibilities:

  • Analyze existing BMC TrueSight/Helix monitoring policies, alerts, thresholds, and event rules
  • Re-map legacy monitoring logic into Datadog and supporting observability pipelines
  • Design and configure log, metric, event, and alert pipelines
  • Support SNMP trap processing, normalization, enrichment, and routing
  • Stand up and configure net-new Datadog-ServiceNow integration
  • Define event correlation, deduplication, escalation, and incident creation logic
  • Support migration from legacy monitoring platforms to Datadog and OpenTelemetry-based pipelines
  • Help design a tiered observability model for production, non-production, and lower-tier workloads
  • Use open-source/OpenTelemetry pipelines where appropriate to optimize Datadog usage and cost
  • Collaborate with infrastructure, network, application, and ITSM teams to align monitoring coverage with operational needs
  • Document monitoring standards, event rules, integration patterns, and operational procedures
  • Support testing, validation, and tuning of monitoring rules and event flows

Required qualifications:

  • Bachelor’s degree in Information Technology, Computer Science, Engineering, or related field, or equivalent practical experience
  • 5–7 years of experience in monitoring, observability, event management, or IT operations engineering
  • Hands-on experience with Datadog or similar observability platforms
  • Experience with BMC TrueSight, BMC Helix, or comparable enterprise monitoring/event management tools
  • Strong understanding of logs, metrics, alerts, events, thresholds, and incident workflows
  • Experience designing or supporting monitoring/event pipelines
  • Experience with SNMP traps, network/device events, and event normalization
  • Familiarity with ITSM platforms, especially ServiceNow
  • Understanding of incident management, alert routing, escalation, and operational support models
  • Ability to analyze existing monitoring rules and translate them into a modern observability framework
  • Strong troubleshooting, documentation, and communication skills

Preferred qualifications:

  • Experience building an event management capability from the ground up
  • Experience with OpenTelemetry collectors, agents, or open-source observability pipelines
  • Experience with Datadog-ServiceNow integration
  • Familiarity with cost optimization strategies for observability platforms
  • Experience supporting enterprise infrastructure, network, cloud, or hybrid environments
  • Knowledge of event correlation, alert noise reduction, and monitoring rationalization
  • Experience working in transformation, migration, or platform modernization programs

What success looks like in this role:

  • Legacy BMC TrueSight/Helix monitoring policies are clearly analyzed, documented, and mapped to target-state observability platforms
  • Datadog is configured to support production-critical monitoring with reliable event and incident flows
  • ServiceNow receives clean, actionable, and properly routed events from Datadog
  • SNMP trap processing is stable, normalized, and aligned with operational needs
  • Lower-tier and non-production workloads are monitored through cost-effective OpenTelemetry-based pipelines
  • Alert noise is reduced, monitoring coverage is improved, and support teams have clearer visibility into incidents
  • Observability standards and documentation are in place for future operational support

What we offer:

  • Competitive pay and salary growth based on your performance.
  • Remote setup, B2B long-term contract, US EST working hours
  • Dynamic working culture with opportunities to grow.
  • Challenging projects with cutting-edge AI and cloud technologies.
  • 100% paid sick leave.
  • Friendly colleagues and a pleasant working atmosphere.

Thank you for your interest in this position. Please note that only candidates whose qualifications closely match our requirements will be contacted.

Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.

Похожие вакансии

Смотреть все →
SREROL-002

Analista SRE Azure

Stefanini Brasil

Sao Paulo

Зарплата не указана

10/08/2026