Platform Engineer
Senior Site Reliability Engineer (SRE)
Удалённо
Откликнуться на сайте работодателяОписание вакансии
We are looking for a
Senior Site Reliability Engineer
to work hands-on with a Grafana-based observability stack and AWS/Kubernetes (EKS), defining SLIs/SLOs, reducing alert noise, building actionable dashboards, strengthening incident response, and improving release safety through progressive delivery and automated deployment analysis.
This is a fully remote position that offers you the flexibility to work from any location in Armenia, whether it's your home or well-equipped offices in Yerevan or Gyumri.
Responsibilities
- Own the observability charter for the platform: build monitoring, alerting, synthetic checks, dashboards, and runbooks
- Define meaningful SLIs/SLOs and reduce alert noise to improve signal quality
- Design and optimize release pipelines with progressive delivery, health gates, and automated rollback mechanisms
- Apply a performance engineering mindset through load testing, capacity analysis, and latency profiling
- Automate operational toil through scripting and infrastructure-as-code
- Accelerate SRE maturity by applying AIOps capabilities to improve detection, diagnosis, and reduce manual effort
- Lead incident response practices including on-call readiness and blameless post-mortems
- Collaborate with DevOps, Cloud teams, product engineering teams, and Tech Leads to drive reliability improvements
Requirements
- 3+ years of experience in Site Reliability Engineering, DevOps, or platform/production engineering supporting customer-facing systems
- Expertise in observability tools such as Grafana, Prometheus, and log/trace aggregation (Loki, Tempo, OpenTelemetry) covering metrics, logs, traces, and events
- Knowledge of SRE fundamentals: SLIs/SLOs, error budgets, golden signals, alert tuning and noise reduction, and blameless post-incident reviews
- Experience operating workloads on Kubernetes (ideally EKS) and AWS, with the ability to debug issues across application, container, and infrastructure layers
- Proficiency in Python, Bash, or Go, with exposure to infrastructure-as-code (Terraform) and CI/CD pipelines
- Background in incident management including triage, escalation, communication, post-mortems, and on-call processes/rotations
- A proactive ownership mindset with the ability to identify problems from telemetry before they're reported and follow through with engineering teams
- Strong communication skills to turn noisy signals into crisp findings, runbooks, and recommendations
- English Level: B2+ (Upper-Intermediate) or higher
Nice to have
- Experience setting up synthetic monitoring (API and browser checks) to validate critical user journeys and catch failures proactively
- Skills in performance engineering, including load/stress testing (k6, JMeter, Locust), capacity planning, and profiling latency, throughput, and resource bottlenecks
- Familiarity with leveraging AIOps capabilities to advance SRE maturity and drive innovation
- Experience with Datadog or similar enterprise observability platforms
- Background in evangelizing best practices and setting standards across engineering teams
- Exposure to programmatic advertising or adtech platforms
We offer
- We connect like-minded people
- Delivering innovative solutions to industry leaders, making a global impact
- Enjoyable working environment, whether it is the vibrant office or the comfort of your home
- Opportunity to work abroad for up to two months per year
- Relocation opportunities within our offices in 55+ countries
- Corporate and social events
- We invest in your growth
- Leadership development, career advising, soft skills and well-being programs
- Certifications, including GCP, Azure and AWS
- Unlimited access to EPAM's internal learning database
- Free English classes with certified teachers
- We cover it all
- Participation in the Employee Stock Purchase Plan
- Monetary bonuses for engaging in the referral program
- Comprehensive medical & family care package
- Four trust days per year for personal needs
- Discounts for fitness clubs
- Benefits package (hotels, restaurants, stores and services)
EPAM is global leader in AI transformation engineering and integrated consulting, serving Forbes Global 2000 companies and ambitious startups. With over thirty years of expertise in custom software, product and platform engineering, we empower our clients to become AI-Native enterprises, driving measurable value from innovation and digital investments.
Это сохранённая копия объявления, опубликованного в другом месте. Вакансии снимают без предупреждения — перед откликом проверьте сайт работодателя. mentors.coach не является нанимающей стороной.