Unlock Full Resume Report

New offer - be the first one to apply!

September 21, 2026

Senior AI Infrastructure Engineer (GPU/HPC, Kubernetes)

Senior • Hybrid

Prague, Czech Republic

Quick Facts

  • Role: Senior AI Infrastructure Engineer (GPU/HPC, Kubernetes)

Description

You will design, operate, and automate the foundational layer of an AI cloud focused on GPU/HPC infrastructure, Kubernetes platforms, CI/CD, Infrastructure as Code, observability, and security. You will help tune performance, plan capacity, scale compute clusters, implement monitoring and alerting, and contribute to secure operations including Confidential Computing. You will also define architecture/operational standards and support critical infrastructure via on-call.

Responsibilities

  • Design, operate, and automate cloud infrastructure for AI and HPC
  • Operate GPU servers, clusters, and Kubernetes platforms for AI workloads
  • Build CI/CD pipelines and Infrastructure as Code
  • Perform performance tuning, capacity planning, and scaling of compute clusters
  • Implement monitoring, dashboards, logs, and alerting for critical components
  • Contribute to secure operations, including Confidential Computing
  • Create architecture and operational standards and mentor teammates
  • Participate in on-call rotations for critical AI infrastructure components

Requirements

  • 5+ years of relevant experience with the ability to take ownership
  • Enterprise Linux administration and automation experience
  • Production Kubernetes including GPU-accelerated workloads
  • Terraform, Ansible, or comparable IaC tools
  • Docker and experience operating container platforms
  • Networking knowledge for high-performance environments and basic server network design
  • Practical experience with GPU servers or clusters
  • Knowledge of the NVIDIA ecosystem (drivers, CUDA/cuDNN/NCCL or NVIDIA Container Toolkit)
  • Monitoring and troubleshooting experience for distributed systems; production support and/or on-call

Benefits

  • Work with modern technologies and large infrastructure
  • Opportunity to technically lead the AI infrastructure and GPU/HPC platform area and set standards
  • Ability to influence solutions from the beginning
  • Small, experienced team with low bureaucracy and room for initiative
  • Hybrid work (home/office) with flexible working hours
  • Indefinite contract and transparent bonus scheme
  • 5 weeks vacation + 5 MyDays, cafeteria benefits, meal allowance, parking, and company recreational facilities
  • Office location in Strahov (near Ladronka Park)

Similar jobs you might like