Unlock Full Resume Report

New offer - be the first one to apply!

September 4, 2026

Senior Site Reliability Engineer

Senior • Remote

Kraków, Poland

Join the Inference Cloud team to design, implement, deploy, and operate AI platforms that enable customers to run inference models and developers to create AI applications with strong performance, compliance, and economics.

Responsibilities

  • Own reliability workstreams for a serverless inference platform.
  • Build automation and tooling, and contribute to architecture and operational decisions.
  • Take ownership of critical reliability problems end-to-end, partner with product engineering teams, and develop expertise in GPU infrastructure, Kubernetes at scale, and AI inference workloads.
  • Build and maintain observability for AI workloads, including telemetry, dashboards, alerts, SLO/SLI tracking, and improvements when targets are missed.
  • Write automation and tooling to reduce operational toil, improve deployment safety, and accelerate incident response.
  • Integrate AI workloads into incident-management processes; build runbooks, participate in on-call rotations, and conduct blameless post-mortems.
  • Build and maintain CI/CD integrations, deployment safety checks, and rollback automation.
  • Collaborate with product engineering teams to improve reliability, contribute to architecture decisions, and ensure operational readiness for product releases.
  • Contribute to capacity planning, autoscaling configuration, and workload scheduling for AI compute infrastructure.

Qualifications

  • Expertise in SRE, infrastructure, or platform engineering, including extensive operational experience managing large-scale distributed systems.
  • Expertise in Kubernetes and large-scale containerization systems.
  • Ability to define SLOs and use observability tools such as Prometheus, Grafana, and distributed tracing to improve system monitoring.
  • Proficiency in Python or Go for automation, CI/CD pipelines, deployment safety, and infrastructure-as-code such as Terraform.
  • Interest in or experience with AI/ML infrastructure, model serving, or GPU workloads.
  • Ability to resolve issues independently while maintaining accountability.
  • Accountability for reliability, automation and monitoring development, and effective collaboration with engineering teams unfamiliar with SRE practices.

Career Development

Development opportunities include programs such as GROW and Mentoring, internal events such as the APEX Expo, and tools such as LinkedIn Learning to expand knowledge and experience.

Similar jobs you might like