Unlock Full Resume Report

New offer - be the first one to apply!

September 16, 2026

Lead Site Reliability Engineer

Senior • Remote

Warsaw, Poland

Quick Facts

  • Role: Lead Site Reliability Engineer

  • Focus: Building SRE reliability practices for an Azure + Kubernetes microservices platform

  • Collaboration: Cloud Engineering, Infrastructure, Development, and management across time zones

Description

You will lead the development of Site Reliability Engineering practices and increase the reliability of a production cloud platform built on Azure, Kubernetes, and microservices. The work centers on implementing SLI, SLO, and error budgets, improving observability, and enhancing incident management processes (including incident escalation support). The role includes autonomy to set SRE standards and define the reliability direction for the organization.

Responsibilities

  • Develop SRE practices and reliability strategy

  • Define and implement SLI, SLO, and error budgets

  • Design centralized dashboards for availability, performance, and system health

  • Expand monitoring, alerting, and observability in production environments

  • Analyze trends and identify improvement areas

  • Support incident management and handle escalations requiring advanced expertise

  • Create and refine operational procedures and RCA standards

  • Collaborate with Cloud Engineering, Infrastructure, and Development on SRE initiatives

  • Help teams identify and eliminate root causes of failures

  • Set SRE roadmap direction and measure initiative progress

  • Prepare business justifications for reliability and infrastructure investments

  • Present recommendations, priorities, and results to senior management

  • Work with a 24/7 incident support function

Requirements

  • Minimum 3 years of experience in Site Reliability Engineering

  • Practical experience implementing SLI, SLO, and error budgets at an organizational level

  • Strong Microsoft Azure knowledge (PaaS and IaaS)

  • Experience with Kubernetes and complex production environments

  • Knowledge of distributed systems, microservices, and cloud-native architecture

  • Experience designing centralized dashboards for performance and availability

  • Knowledge of Prometheus, Grafana, Elasticsearch, and Elastic APM

  • Experience monitoring complex cloud infrastructure

  • Knowledge of incident management processes; ability to create SOPs and perform RCA

  • Ability to drive cross-team technical initiatives

  • Ability to collaborate across different organizational levels

  • Ability to communicate technical topics in terms of business goals, risks, and value

  • Experience creating business justifications, recommendations, and management presentations

  • Very good English in speech and writing

  • Independence, analytical mindset, and ability to solve complex problems

Benefits

  • Support for professional and personal development, aligned with your goals

  • Individual support from a Service Delivery Manager

  • Training, certificates, and conference costs supported or fully covered

  • #SmartChange: opportunity to change projects based on your preferences

  • Work-life balance activities and integration events

  • Physical activity support and access to training spaces

  • Health package: private care, fitness card, insurance, and psychological support platform

  • Flexible benefits: choose how to allocate your benefit points

Similar jobs you might like