Unlock Full Resume Report

New offer - be the first one to apply!

September 17, 2026

Platform Reliability Engineer

Senior • Hybrid • On-site

Prague, Czechia

Quick Facts

  • Role: Platform Reliability Engineer

  • Focus: Monitoring, incidents, and alert routing; sustainable reliability improvements (no on-call)

Description

Join the team to strengthen how systems are monitored, how incidents are handled, and how alerts are routed—so engineering teams can ship with confidence. You’ll operate and improve the monitoring/observability stack (Prometheus, Grafana, OpenTelemetry), help define incident processes and post-incident learning artifacts, and work with engineers to make reliability standards practical.

Responsibilities

  • Operate and improve the monitoring stack: instrument services, define what to watch in production, and shape alerting for actionable signals without noise.

  • Help define incident operations: clear communication, structured learning afterward, and supporting artifacts such as a status page and runbooks.

  • Collaborate with platform and product engineers to adopt better reliability tooling and practices as systems change; write documentation teams actually use.

Requirements

  • Hands-on experience choosing what to measure in production based on real customer experience signals.

  • Comfortable with incidents and alerts from early detection through resolution and follow-up to prevent recurrence.

  • Hands-on experience with Prometheus, Grafana, OpenTelemetry (or similar).

  • Hands-on experience with alert-routing tools such as PagerDuty.

  • Ability to read and write code and follow services/pipelines across the stack.

  • Ability to support a blame-free, learning-focused post-incident culture.

  • Ability to write clear, concise guidance that teams adopt.

  • Driven to automate repetitive tasks and improve developer workflows.

Benefits

  • Full-time role with space, support, and autonomy for personal growth and direct impact.

  • Offices in Prague (100Yards) or Brno (Titanium), with option to work remotely.

  • Flexible working hours.

  • Unlimited Claude for every team member.

  • Stock options and profit sharing.

  • Solid education and training budget, conference tickets, and internal learning sessions.

  • Generous hardware budget; free lunches when in the office.

  • Unlimited coffee/beer and snacks.

  • Free entry to Prague Zoo; free Multisport card.

  • Team events/offsites and various office activities.

First 3 Months Expectations

  • Complete onboarding and align on collaboration/production-response responsibilities.

  • Understand the platform at a high level and handle smaller infrastructure problems, incidents, or bugs.

  • Map the current monitoring, incidents, and alerts process to identify friction and improvement opportunities.

  • Publish initial monitoring/observability/alerting guidelines (signals, naming, dashboards, alerting principles).

  • Participate in incident reviews and turn patterns into improved playbooks.

  • Contribute to team ceremonies and technical discussions to support infrastructure needs.

First 6 Months Expectations

  • Work mostly independently on larger tasks, while staying able to ask for help.

  • Build cross-team relationships and gather feedback to improve daily engineering workflows.

  • Have teams reference your guidance for higher-risk changes with measurably less alert noise and duplicate paging.

  • Maintain incident documentation that is easy to find and actually used.

  • Own the monitoring and alerting improvement roadmap end-to-end; align priorities with leadership.

Similar jobs you might like