Unlock Full Resume Report

New offer - be the first one to apply!

September 17, 2026

Senior Site Reliability Engineer

Senior • On-site

Warsaw, Poland

Quick Facts

  • Role: Senior Site Reliability Engineer

Description

Lead and coordinate team-level projects to improve system reliability across a massive distributed architecture. Drive the design and implementation of complex reliability solutions and internal platforms, including incident response automation with an AI agent, and run major outages as a senior Incident Commander with blameless post-mortems. Build and scale self-service performance testing and automated resiliency platforms (fault injection and data center failover experiments) while owning reliability, managing technical debt, and automating infrastructure.

Responsibilities

  • Coordinate reliability improvement projects across a distributed microservices environment

  • Design and implement reliability platforms and tools

  • Build an AI agent to enhance incident response workflows

  • Serve as Incident Commander during major outages and lead blameless post-mortems

  • Develop and scale self-service performance testing with real-time observability

  • Build automated resiliency via continuous fault injection and lead failover experiments

  • Own team systems reliability, manage technical debt, and automate infrastructure

  • Mentor engineers and set code quality standards

  • Influence architecture decisions and team communication on complex problems

Requirements

  • Solid engineering background with deep understanding of distributed systems and microservices architecture

  • Proficient in at least one of: Kotlin, Java, Python, Go

  • Hands-on advanced container/orchestration experience: Kubernetes and Docker, plus networking and service mesh concepts

  • Hands-on experience with Chaos Engineering (fault injection, failure scenarios)

  • Practical knowledge of Gatling or similar performance testing tools is a plus

  • Deep experience with monitoring/observability stacks such as Prometheus, Grafana, and ELK

  • Ability to identify single points of failure and act as Incident Commander

  • Strong practical knowledge of CI/CD pipelines and Infrastructure as Code (IaC)

  • Ownership mindset, data-driven approach, and mentoring experience

Benefits

  • Flexible working hours in a hybrid model (4/1), with occasional remote work (30 days)

  • Long-term discretionary incentive plan based on Allegro.eu restricted stock units

  • Annual bonus based on individual performance and company results

  • Well-located offices with excellent work tools

  • MacBook Pro or Dell equivalent with Windows and accessories

  • Fringe benefits via a cafeteria plan (medical, sports/lunch packages, insurance, vouchers)

  • English classes related to job responsibilities

  • Training budget, inter-team opportunities, hackathons, internal learning platform

  • Additional day off for volunteering

  • Social events (e.g., Spin Kilometers, Family Day, Fat Thursday, Advent of Code)

Similar jobs you might like