Unlock Full Resume Report
ATS Pass
Missing keywords
Tailored AI suggestions
Job match analysis
Interview-focused insights
New offer - be the first one to apply!
September 17, 2026
Platform Reliability Engineer
Senior • Hybrid • On-site
Prague, Czechia
Apply now
Quick Facts
Role: Platform Reliability Engineer
Focus: Monitoring, incidents, and alert routing; sustainable reliability improvements (no on-call)
Description
Join the team to strengthen how systems are monitored, how incidents are handled, and how alerts are routed—so engineering teams can ship with confidence. You’ll operate and improve the monitoring/observability stack (Prometheus, Grafana, OpenTelemetry), help define incident processes and post-incident learning artifacts, and work with engineers to make reliability standards practical.
Responsibilities
Operate and improve the monitoring stack: instrument services, define what to watch in production, and shape alerting for actionable signals without noise.
Help define incident operations: clear communication, structured learning afterward, and supporting artifacts such as a status page and runbooks.
Collaborate with platform and product engineers to adopt better reliability tooling and practices as systems change; write documentation teams actually use.
Requirements
Hands-on experience choosing what to measure in production based on real customer experience signals.
Comfortable with incidents and alerts from early detection through resolution and follow-up to prevent recurrence.
Hands-on experience with Prometheus, Grafana, OpenTelemetry (or similar).
Hands-on experience with alert-routing tools such as PagerDuty.
Ability to read and write code and follow services/pipelines across the stack.
Ability to support a blame-free, learning-focused post-incident culture.
Ability to write clear, concise guidance that teams adopt.
Driven to automate repetitive tasks and improve developer workflows.
Benefits
Full-time role with space, support, and autonomy for personal growth and direct impact.
Offices in Prague (100Yards) or Brno (Titanium), with option to work remotely.
Flexible working hours.
Unlimited Claude for every team member.
Stock options and profit sharing.
Solid education and training budget, conference tickets, and internal learning sessions.
Generous hardware budget; free lunches when in the office.
Unlimited coffee/beer and snacks.
Free entry to Prague Zoo; free Multisport card.
Team events/offsites and various office activities.
First 3 Months Expectations
Complete onboarding and align on collaboration/production-response responsibilities.
Understand the platform at a high level and handle smaller infrastructure problems, incidents, or bugs.
Map the current monitoring, incidents, and alerts process to identify friction and improvement opportunities.
Publish initial monitoring/observability/alerting guidelines (signals, naming, dashboards, alerting principles).
Participate in incident reviews and turn patterns into improved playbooks.
Contribute to team ceremonies and technical discussions to support infrastructure needs.
First 6 Months Expectations
Work mostly independently on larger tasks, while staying able to ask for help.
Build cross-team relationships and gather feedback to improve daily engineering workflows.
Have teams reference your guidance for higher-risk changes with measurably less alert noise and duplicate paging.
Maintain incident documentation that is easy to find and actually used.
Own the monitoring and alerting improvement roadmap end-to-end; align priorities with leadership.
Similar jobs you might like

Senior Site Reliability Engineer (AI/ML Platform)
Link Group
Senior
Technology
Warsaw, MZ, Poland · Remote
20K zł - 27K zł/yr
12 days ago

Site Reliability Engineer (SRE)
Yard Corporate
Senior
Technology
Warsaw, Poland · On-site
40K zł - 55K zł/yr
20h ago

Site Reliability Engineer (SRE)
EPAM Systems
Mid
Technology
Lodz, ŁD, Poland · Remote
N/A
20h ago
Fullstack Support Engineer (Java/React)
Grape Up
Mid
Technology
Krakow, Poland · Remote
17K zł - 22K zł/yr
8 days ago

Run Engineer N3 /Site Reliability Engineer (h/f)
emagine Polska
Senior
Technology
Paris, France · On-site
N/A
13 days ago
Lead Site Reliability Engineer
Edge One Solutions Sp. z o.o
Senior
Technology
Warsaw, Poland · Remote
N/A
20h ago

Site Reliability Engineer
Link Group
Senior
Technology
Warsaw, Poland · On-site
N/A
20h ago

Senior Site Reliability Engineer (AI Hardware & Infrastructure)
Link Group
Senior
Technology
Warsaw, MZ, Poland · Remote
20K zł - 27K zł/yr
12 days ago

Engineering Team Lead - Platform Engineering
RTB House
Senior
Technology
Warsaw, Poland · Remote
0K zł - 0K zł/yr
20h ago
Site Reliability Engineer, Senior
Webellian Sp.z o o
Senior
Technology
Warsaw, MZ, Poland · On-site
N/A
11 days ago