Unlock Full Resume Report

New offer - be the first one to apply!

September 4, 2026

Senior Site Reliability Engineer

Senior • On-site

Warsaw, Poland

We are seeking a Senior Site Reliability Engineer to own the reliability, observability, and operational health of production AI systems. The role bridges deployment and long-term operability while embedding cost, security, and quality discipline throughout each solution's lifecycle.

Responsibilities

  • Own end-to-end deployment, including infrastructure as code, CI/CD pipelines, and Azure environment management; ensure every environment can be rebuilt from source.
  • Build LLM-aware observability with traces for every model call, production quality signals such as evaluation sampling, drift detection, and guardrail-trigger rates, plus cost and latency dashboards.
  • Define and defend SLOs for availability, latency, and quality objectives per solution; balance delivery speed and stability using data-driven error budgets.
  • Run incident management, including on-call models, pre-written runbooks, and blameless postmortems.
  • Manage AI workload costs by monitoring token economics per solution, implementing budgets and alerts, and conducting proactive capacity planning.
  • Maintain security through patching, secret rotation, access reviews, and audit readiness across the full solution lifecycle.
  • Shape operability requirements before handover, incorporate them into the pod's definition of done, and run joint hypercare with operational sign-off.
  • Feed operational patterns, failure modes, and cost learnings back to pods and the Architect.

Requirements

  • 5+ years of experience operating cloud production systems, with a track record of scaling, defining SLOs, managing on-call rotations, and automating manual work.
  • Expertise in Azure IaaS/PaaS operations, infrastructure as code using Terraform or Bicep, and CI/CD tooling.
  • Proficiency with observability stacks, LLM tracing, container orchestration, and Python/Bash automation.
  • Knowledge of FinOps fundamentals for AI workloads.
  • Familiarity with daily AI use in operations, including incident triage, runbook drafting, log analysis, and automation-code development.

We offer

  • Hybrid-by-design work model and the opportunity to work remotely within Poland.
  • Opportunity to work abroad for up to 60 days annually.
  • Business-driven relocation opportunities.
  • Career development programs, thought leadership, mentoring, soft-skills and well-being programs.
  • Certification opportunities in Anthropic, Gemini, GCP, Azure, and AWS.
  • English classes.
  • Stable pay.
  • Participation in the Employee Stock Purchase Plan with a 15% discount.
  • Benefits package including health insurance, multisport, and shopping vouchers.
  • Referral bonuses up to $2,000.
  • Offices with entertainment and relaxation zones, table tennis, football, free snacks, coffee, and more.
  • Corporate, social, and well-being events.

Benefits listed above are available to employees only.

Contractor cooperation is available; B2B agreement terms are agreed individually.

Only selected candidates will be contacted.

Similar jobs you might like