New offer - be the first one to apply!

October 1, 2026

Senior Site Reliability Engineer

Senior • Remote

Kraków, Poland

Quick Facts

  • Role: Senior Site Reliability Engineer

Description

As a Senior SRE, you will own reliability workstreams for a serverless AI inference platform, including building observability, automation, and tooling to improve deployment safety and incident response. You will partner with product engineering teams to shape architecture and operational decisions, contribute to runbooks and on-call rotations, and conduct blameless post-mortems. You will also support capacity planning, autoscaling, and workload scheduling for AI compute infrastructure, with opportunities to deepen expertise in GPU infrastructure and Kubernetes at scale.

Responsibilities

  • Build and maintain observability for AI workloads: telemetry, dashboards, alerts, and SLO/SLI tracking; drive improvements when targets are missed

  • Develop automation and tooling to reduce operational toil, improve deployment safety, and accelerate incident response

  • Integrate AI workloads into incident management processes: runbooks, on-call rotations, and blameless post-mortems

  • Build and maintain CI/CD integrations, deployment safety checks, and rollback automation

  • Collaborate with product engineering to improve reliability, contribute to architecture decisions, and ensure operational readiness for releases

  • Contribute to capacity planning, autoscaling configuration, and workload scheduling for AI compute infrastructure

Requirements

  • Expertise in SRE, infrastructure, or platform engineering managing large-scale distributed systems with extensive operational experience

  • Expertise in Kubernetes and large-scale containerization systems

  • Ability to define SLOs and use observability tools such as Prometheus, Grafana, and distributed tracing

  • Proficiency in Python or Go for automation, CI/CD pipelines, deployment safety, and infrastructure-as-code with Terraform

  • Interest in or experience with AI/ML infrastructure, model serving, or GPU workloads

  • Ability to resolve issues independently while maintaining accountability throughout the process

  • Accountability for reliability through automation and monitoring, and effective collaboration with product engineers unfamiliar with SRE practices

Benefits

Flexible benefit options designed to meet individual needs, including coverage for health, finances, family, and time at work and for other endeavors.

Similar jobs you might like

Unlock Full Resume Report