Unlock Full Resume Report

New offer - be the first one to apply!

September 4, 2026

Senior Site Reliability Engineer (AI Hardware & Infrastructure)

Senior • Remote

20,000 - 27,000 PLN/yr

Warsaw, MZ, Poland

We are looking for an elite Site Reliability Engineer to join the core team responsible for global AI compute infrastructure. The mission is to ensure that the physical and virtualized backbone of the AI platform—bare-metal servers, high-density GPU racks, and the advanced network that connects them—is reliable, performant, and scalable.

This is a hands-on engineering role requiring comfort with Python automation and Infrastructure-as-Code, BGP routing strategies, and collaboration with data center technicians.

Quick Facts

  • Support global AI compute infrastructure, including bare-metal servers, high-density GPU racks, virtualized hardware, and advanced networking.
  • Build automation and operational tooling at fleet scale.
  • Participate in a 24/7 on-call rotation and lead responses to high-severity incidents.

What You Will Do

  • Write sophisticated Python tooling and automation to manage the complete server-fleet lifecycle, including provisioning, configuration, maintenance, and decommissioning across thousands of machines.
  • Integrate operational systems such as JIRA, Siebel, and PagerDuty through robust APIs to automate workflows and reduce hardware and network incident resolution time.
  • Design and implement an observability stack for bare-metal and virtualized hardware, including custom telemetry pipelines, Grafana dashboards using Prometheus data, and AI-driven anomaly detection to predict failures.
  • Lead critical-incident response, participate in the 24/7 on-call rotation, spearhead high-severity outage response, and drive blameless post-mortems that produce architectural improvements.
  • Use advanced AI utilities and LLM-assisted development to improve technical execution, including automation-script generation and system-performance evaluation.

Who We’re Looking For

  • Deep Site Reliability or Production Engineering background, supported by a solid Computer Science foundation and experience managing large-scale, mission-critical infrastructure.
  • Exceptional Python programming ability, with experience building scalable, robust operational tools and automation frameworks from the ground up.
  • Practical expertise in advanced network topologies, high-bandwidth routing and switching, BGP, and dual-stack IPv4/IPv6 environments.
  • Expert hands-on experience with Prometheus, Grafana, OpenTelemetry, and Loki.
  • Extensive experience designing service rollout strategies, defining meaningful alert thresholds, creating technical runbooks, and leading incident-response war rooms.
  • Ability to own ambiguous, complex technical problems from initial concept through a production-grade, fully automated solution.
  • Ability to collaborate with external data center vendors and on-site field technicians to maximize uptime.

Similar jobs you might like