Skip to content
Site Reliability Engineer

Site Reliability Engineer

Customer Service Reliability Engineer

Thales

Praguefulltimesenior

Apply on the employer's site

Role description

Location: Praha, Czechia
Thales people architect identity management and data protection solutions at the heart of digital security. Business and governments rely on us to bring trust to the billions of digital interactions they have with people. Our technologies and services help banks exchange funds, people cross borders, energy become smarter and much more. More than 30,000 organizations already rely on us to verify the identities of people and things, grant access to digital services, analyze vast quantities of information and encrypt data to make the connected world more secure.
Thales in the Czech Republic employs over 400 people from 45 different nationalities. A total of 15 teams work on projects for government agencies, banking, mobile services and the Internet Of Things (IoT) technology. At the core of our business is the development of software which we configure and embed in a multitude of different devices and form factors. These include many kinds of payment cards, SIM cards, travel passes, secure eBanking devices, authentication tokens, machine identification modules (MIM), and secure ID documents including ePassports, eID and eHealth cards, as well as eDriving licenses. Because of the international environment surrounding us every day, it comes as no surprise that English is our official corporate language.
Thales is looking for a Customer Reliability Engineer (CRE) who is primarily responsible to ensure the best customer experience by assuring services reliability from customer eyes and making sure Incident/Service Requests are resolved in the shortest timeframe.

In this position, you will also be responsible to ensure overall service quality and health by closing the loop with feedbacks to Product Owner & Site Reliability Engineering regarding service issues and improvements. You will improve service adoption and customer success by ensuring customers are getting the most from our Cloud Services.

Key Areas of Responsibility

  • Manage Incidents/Service requests within SLA, tracking SLA and taking actions in case of deviation.
  • Escalate to SRE or Engineering (L3) Incidents/Service requests that cannot be resolved.
  • Communicate with customers by keeping them informed about the updates related to Incidents/Service requests on a regular basis.
  • Write Work Instructions/Incident Response Plans for L1, automate WI based on alert.
  • Write Technical Notes/Knowledge Articles for CRE and Customers (focused on service usage improvement).
  • Maintain high technical skills on solution/services, be Subject Matter Expert (SME) of selected solution/service.
  • Build and deliver technical webinar to Customers and other CRE, work with Solution Designer/Product Owner to build and Customer Success Manager/Customer Reliability Service Manager to plan such webinar.
  • Be Customer Champion for selected accounts, establish privileged relationship and deeper technical understanding for selected accounts, stay up-to-date on their plans regarding service usage.
  • Deploy Customer's specific changes upon Change Approval Board approval.
  • Implement and maintain Service Level Indicator, dashboards and Customer's specific alerts to follow performance/improvement plan.
  • Follow-up Customer activity through dashboards (percentage of success, percentage of enrollment, percentage of conversion, etc.).
  • Provide close support to Customer Reliability Service Manager when it comes to understanding of customer use cases and Incidents/Service requests.
  • Scale up/down to meet customer business need (if possible at CRE level, or raise the need to SRE).
  • Lead Root Cause Analysis when no SRE involved (if there's a service outage, SRE will be involved and will naturally become the RCA leader).
  • Participate in post-mortems and contribute to RCA.
  • Translate internal RCA to external RCA, publish external RCA in due time (according to service/customer agreement).
  • Review repeated incidents or known error with Product Owner/Service Reliability Engineer.
  • Raise product/service improvement requests to PO/SRE.
  • CRE is working on-call to provide 365x24x7 upon L1 escalation.

Minimum Requirements

  • Bachelor’s Degree in Computer Science, Software Engineering, or equivalent degree
  • Intermediate-Advanced English fluency
  • At least 5 years of experience as CRE/SRE.
  • Mandatory experience with Cloud environments like AWS or GCP
  • Java is preferred.
  • Notions of Databases (Mongo DB) and SQL queries (MS-SQL Server and Oracle).
  • Knowledge in a few development languages front-end, Angular, Javascript, XAML, and styling CSS, bootstrap, Material Design.
  • Mandatory experience in Linux, and a good understanding of IT security principles (PKI).
  • It would be preferred if you have experience with Shell Scripting, Python, Terraform.
  • Knowledge and experience with Splunk, Confluence/Jira, Snow is good to have.
  • Strong team player with proven experience and a willingness to take ownership of a topic and successfully bring it to completion.
  • Well organized with strong attention to detail, strong verbal and written communication skills.

At Thales we provide CAREERS and not only jobs. With Thales employing 80,000 employees in 68 countries our mobility policy enables thousands of employees each year to develop their careers at home and abroad, in their existing areas of expertise or by branching out into new fields. Together we believe that embracing flexibility is a smarter way of working. Great journeys start here, apply now!

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

Coming to this page

A resume for this role — and a ticket to the draw

We take the posting apart down to the real requirements and rewrite your resume against it — by asking, not inventing: no line appears without your confirmation. Sign in to get it first, and to enter the draw.

  • A resume for this exact role, not a universal one
  • Answers are kept: edit any one, not the whole conversation
  • All in your account — open it from any device

On the wheel

A discount on mentoring

Winners are drawn at random among entries with a confirmed email. The date and the full rules are on the draw page.

Draw rules

Site Reliability Engineer

Software Engineer

Microsoft

Praguefulltime

Apply on the employer's site

Role description

Overview

We are seeking a skilled software engineer to join our team and help implement advanced Identity and Access Management standards by leveraging emerging AI-forward technologies. In this role, you will work on complex, high‑impact technical challenges in close collaboration with subject matter experts, engineers, and architects across Substrate, Microsoft 365, E+D, Entra, and Azure. These initiatives offer meaningful opportunities for deep technical growth and long‑term career progression.The ideal candidate is passionate about building scalable, secure solutions for a broad set of customers, including service developers, and consistently delivers high‑quality systems aligned with industry best practices in security and reliability. You bring strong problem‑solving and debugging skills, along with a solid foundation in modern software engineering practices, including SDK and shared component development for hyperscale distributed systems.Success in this role requires a strong sense of system design, a continuous improvement mindset, and an uncompromising focus on quality.Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.

Responsibilities

Leads efforts to use and enhance, or build, new software developer tools to support easier, faster, and more effective software engineering across products. Identifies whether open source or internal code is available to address coding needs for a set of products, and reuses it in a responsible manner where applicable. Develops substantial skills in tools inside and outside current areas of expertise. Leads identification and/or creation of tools that are useful for building the product. Shares best practices and teaches others about new tools and strategies.Leads efforts to ensure the correct processes are followed to achieve a high degree of security, privacy, safety, and accessibility across solutions and teams. Creates and assures the presence of visible evidence (e.g., audit trail) to demonstrate compliance for products. Develops and maintains a deep understanding of the implications of onboarding new technologies following expectations of compliance at Microsoft. Demonstrates and maintains an up-to-date understanding of both global and local regulations for technologies and system applications to ensure regulations are followed and met.Understands and applies security best practices and establishes code invariants to model "security as code," ensuring each layer is independently secure, and minimizing risk. Supports and/or adopts, and may set security standards for clear security code review practices for a set of products that align with design and engineering principles to raise the security hardening for both protections and detections. Proactively incorporates deployment gates on security controls, and scanners for a set of products to prevent regressions and/or vulnerabilities that would have customer impact. Includes required security monitoring to ensure detection of violations. Collaborates with relevant security partners to define security promises and security invariants for the design of a product/solution while factoring in attacker/investigator personas for security monitoring and telemetry needs, ensure threat models and premortems validate upstream and downstream assumptions and security invariants, establish security breach drills and security incident response processes (e.g., impact analysis, containment), and ensure that artificial intelligence (AI) safety features are implemented for the AI production systems tied to a set of products.Collaborates with partner teams to ensure a set of products work well with the components of the partner team, ensuring proper end-to-end testing, live-site coverage, scalability, performance, and DRI escalation pathways are established before going live.Considers and leads the identification of requirements for, and the comprehensive application of automation within production and deployment across products, targeting zero-touch deployment when possible. Runs code in simulated or other non-production environments to confirm functionality and error-free runtime across products.ImplementLeads efforts for experiments that determine the impact of changes using feature flags/flighting in their code, interprets results, and decides on next steps or ship decision from results. Drives identification of the correct metrics for experimentation in determining improving customer value. Drives collaboration efforts with internal partners (e.g., Data Science, product managers) to ensure incorporation of success and guard rail metrics for experimentation.Leverages their subject-matter expertise to partner with appropriate stakeholders (e.g., technical program managers) to drive multiple groups' project plans, release plans, and work items. Breaks down long-term project vision into milestones as part of an overall roadmap. Guides other members for project estimation and escalates issues that might cause a major delay. Drives efforts to ensure required security protections and detection processes are accounted for in planning. Drives efforts to ensure project plans adhere to security, privacy, and compliance requirements. Drives efforts to ensure all code for a set of products/solutions is properly flighted for quicker mitigation of production incidents. Calculates capacity for planning, accounting for appropriate failover and backup/restore mechanisms for disaster recovery for a set of products and/or solutions. Makes considerations for efficient operation of a set of products and/or solutions after it is live. Proactively establishes rollback plans for a set of products and/or solutions.Leads leveraging existing deployment frameworks in the implementation of solutions within the existing framework, automating deployment tasks when possible to ensure efficiency. Proactively follows safe change deployment best practices (e.g., ensuring that flights are set correctly) for their team to minimize adverse impact to users and other services. Optimizes deployments within products and components to meet differing business objectives. Leads efforts to ensure that solutions are deployed safely, rolling out security-sensitive features only to applicable, relevant customers and scenarios to reduce the attack surface. Proactively monitors dependency status and ensures that only the latest, secure versions are deployed. Defines when rollback plans should be enacted for a set of products. Drives building deployment infrastructure to allow developers' private builds for a set of products/solutions to be tested in a production-like environment.

Reliabilty and SupportabilityIntegrates, designs, and reviews others' work across a team or product to integrate logging and instrumentation for gathering telemetry data on system behavior such as performance, reliability, availability, usage, and safety mechanisms, and for allowing monitoring and investigating security-related concerns and scenarios for both live and A/B experiments for products, services, and offerings. Leverages telemetry feedback and effectiveness to drives the improvement of subsequent monitoring designs. Ensures solutions are scalable, financially responsible, and meet capture/storage guidelines. Leads efforts to classify, and analyze complex data and analyses on a range of metrics (e.g., health of the system, where bugs might be occurring), and leads the creation of outputs (e.g., notifications, dashboards) that improve monitoring and investigating security-related concerns and scenarios, system monitoring and/or issue identification and mitigation. Proactively considers the privacy implications of telemetry code changes, and of adding new data points.Holds accountability as a designated responsible individual (DRI) and mentors other engineers across products/solutions, working on-call to monitor system/product/service for degradation, downtime, or interruptions. Alerts stakeholders as to status and initiates actions to restore system/product/service for complex issues. Develops a playbook for the team to resolve issues. Coordinates people and resources to ensure DRI responsibilities are covered across teams. Responds within service level agreement (SLA) timeframe. Has line of sight to incidences and plans to address emerging issues. Leads efforts to reduce incident volume, looking globally at incidences and providing broad resolutions. Escalates issues to appropriate owners.Maintains operations of live site service, following security best practices when responding quickly to mitigate issues while using the minimum required permissions to do so that arise on a rotational, on-call basis. Implements and helps others implement solutions and mitigations to complex issues impacting the performance or functionality of live site services. Reviews and writes incident postmortem and presents insights that drive changes to reduce or eliminate incidents. Proactively improves troubleshooting guides (TSGs), wikis, tests, and telemetry to make on-call better, and recommends user-facing support documentation and additional test coverage to reduce likelihood of future user-initiated incidents. Enables secure operations, security monitoring, and integration with live site investigation activities. Proactively identifies opportunities (e.g., lunch talks, automation, practices, tools) that can be leveraged to improve the live site experience and executes on them.

Qualifications **Required Qualifications:**Bachelor's Degree in Computer Science or related technical field AND technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or PythonOR equivalent experience.

**Preferred Qualifications:**Master's Degree in Computer Science or related technical field AND technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or PythonOR Bachelor's Degree in Computer Science or related technical field AND technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or PythonOR equivalent experience.

Software Engineering IC3 - The typical base pay range for this role across Czechia is Kč 966,000.00 - Kč 1,727,000.00 per year. Certain roles may be eligible for benefits and other compensation.

Find additional benefits and pay information here:

https://careers.microsoft.com/v2/global/en/corporate\-pay/czech\-republic\-corporate\-pay.html

This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.

Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process**.**

This is a saved copy of a posting published elsewhere. Postings get taken down without notice — check the employer's site before applying. mentors.coach is not the hiring party.

That is every opening in this category and country

Customer Service Reliability EngineerThales · Czechia

Apply on the employer's site