Unlock Full Resume Report

New offer - be the first one to apply!

September 2, 2026

Site Reliability Engineer

Senior • Remote

17,850 - 22,700 PLN/yr

Warsaw, Poland

We are looking for a Site Reliability Engineer to define and drive the reliability of systems at the scale of millions of clients. In this role, you will strengthen SRE practices and shape the resilience of the technology stack through high-impact observability, ensuring systems remain robust and scalable.

Responsibilities

Observability Platform Engineering

  • Develop a standardized observability ecosystem.
  • Implement a conscious telemetry model focused on structured events, distributed tracing, and intelligent sampling strategies that provides deep, actionable insights into system behavior.

Reliability Enablement

  • Act as a strategic partner to product engineering teams, providing the platform, standards, and data they need to own service reliability.
  • Use error budgets and alerting to balance feature velocity with stability.

Proactive Resilience & Protection

  • Enhance detection capabilities to identify issues before they impact customers.
  • Leverage early-warning systems and AI/ML for automated anomaly detection and intelligent data analysis to continuously verify and strengthen system resilience.

Operations & Tooling

  • Build internal automation and tooling that streamlines SRE workflows, automates routine operational tasks, and enhances efficiency across the technology stack.

Incident Management & On-Call Rotation

  • Participate in an on-call rotation to provide incident management, ensuring rapid incident resolution, effective communication, and post-incident analysis to drive continuous improvement.

Requirements

  • Professional experience in SRE, Infrastructure, or DevOps roles managing high-scale, distributed environments.
  • Advanced programming skills in Python, with a strong focus on building scalable automation, internal tooling, and robust scripts.
  • Hands-on expertise managing production-grade Kubernetes environments, configuration management tools such as Ansible, and designing resilient infrastructure architectures within Azure Kubernetes Service and on-premises environments.
  • Proficiency in building standardized telemetry ecosystems and using self-hosted open-source observability tools for data collection, storage, and visualization, including Prometheus, Grafana, ELK Stack, Tempo, Thanos, Jaeger, and similar tools.
  • Ability to drive incident management, conduct thorough post-incident analysis, and foster a culture of reliability and shared ownership.

Nice to Have

  • Experience with commercial APM platforms such as Datadog, Splunk, or New Relic, and chaos engineering tooling.
  • Experience with cloud cost management and FinOps principles.
  • Experience defining and tracking SRE metrics (SLI/SLOs) and managing error budgets.
  • Experience with AI/ML techniques for SRE tasks, including AIOps, automated anomaly detection, log analysis, and reliability-workflow optimization.
  • Experience building and managing strategies to proactively manage technical debt and align team output with organizational goals.

What We Offer

  • Real influence on the development of the company and product.
  • Work in an experienced team that shares knowledge.
  • Regular feedback and clear career paths.
  • Regular team-building meetings.

Benefits

  • Training budget for courses and conferences.
  • An extra day off on your birthday.
  • An extra day off for parents.
  • Equipment tailored to your needs.
  • Private medical care and group insurance.
  • Access to an e-learning platform for English learning and a benefits platform.
  • Access to a wellbeing platform, workshops, and private therapy sessions.
  • Remote work, work from the Warsaw office, or a coworking space in your city.

Similar jobs you might like