Unlock Full Resume Report
ATS Pass
Missing keywords
Tailored AI suggestions
Job match analysis
Interview-focused insights
New offer - be the first one to apply!
September 22, 2026
Senior Site Reliability Engineer
Senior
205,000 - 290,000 USD/yr
WY
Apply now
Quick Facts
- Senior production engineering role focused on SRE, reliability automation, and Kubernetes-based GPU cluster operations
Description
Build and operate automation for large-scale Kubernetes clusters across cloud partner and on-prem environments. Develop tools and services for provisioning, validation, upgrades, monitoring, repair, and cluster lifecycle operations, while improving Day 0/1/2 workflows and reducing manual production touches through APIs and GitOps. Define SLOs/SLIs, monitor error allowances, streamline reporting, and participate in on-call incident response and durable follow-up.
Responsibilities
- Build and operate automation for large-scale Kubernetes clusters (NCP and on-prem)
- Develop tooling for cluster provisioning, validation, upgrades, monitoring, repair, and lifecycle operations
- Improve Day 0/Day 1/Day 2 workflows for bringup, handoff, and production operations
- Define SLOs/SLIs and monitor error allowances; streamline reliability reporting
- Reduce manual production touches using APIs, GitOps, and agent-assisted workflows
- Participate in on-call, incident response, debugging, and durable follow-up
- Partner with platform, storage, networking, security, and workload teams to make infrastructure production-ready
Requirements
- 8+ years of experience building or operating production infrastructure
- Strong programming skills in Python, Go, or similar
- Expert-level knowledge of Linux, Kubernetes, containers, cloud infrastructure, or infrastructure automation
- Solid grasp of SRE principles: SLOs, SLIs, error budgets, incident management
- Ability to troubleshoot distributed systems in production
- Experience building and operating comprehensive observability stacks (monitoring, logging, tracing), including OpenTelemetry, Prometheus, Grafana, ELK Stack, Lightstep, Splunk
- Clear communication and cross-team collaboration
- BS/MS in Computer Science or equivalent experience
Benefits
- Base salary determined by location and experience; equity eligibility and benefits included
- Base salary range: 168,000 USD - 270,250 USD (Level 4) and 208,000 USD - 333,500 USD (Level 5)
Similar jobs you might like

Site Reliability Engineer
Link Group
Senior
Technology
Warsaw, Poland · On-site
N/A
5 days ago

Senior Engineer - SRE & Infrastructure Services
EPAM Systems
Senior
Technology
Wrocław, Poland · Remote
N/A
5 days ago

Senior Site Reliability Engineer - GIPHY
Shutterstock
Senior
Technology
New York, NY
$150K - $150K/yr
4 days ago
Devops Engineer
RemoDevs
Senior
Technology
Warsaw, Poland · Remote
23K zł - 24K zł/yr
5 days ago

Senior Site Reliability Engineer
Xopero Software
Senior
Technology
Gorzów Wielkopolski, Poland · Remote
240K zł - 332K zł/yr
4 days ago
Lead Site Reliability Engineer | Branża Technologiczna
Edge One Solutions Sp. z o.o
Senior
Technology
Warsaw, Poland · Remote
N/A
5 days ago

Site Reliability Engineer (SRE)
Yard Corporate
Senior
Technology
Warsaw, Poland · On-site
40K zł - 55K zł/yr
5 days ago

Senior Engineer - SRE & Infrastructure Services
EPAM Systems
Senior
Technology
Poznań, Pl-30, Poland · Remote
N/A
5 days ago

Senior Site Reliability Engineer
EPAM Systems
Senior
Technology
Warsaw, Poland · On-site
N/A
17 days ago

Site Reliability Engineer (SRE)
EPAM Systems
Mid
Technology
Lodz, ŁD, Poland · Remote
N/A
5 days ago