Unlock Full Resume Report

New offer - be the first one to apply!

September 16, 2026

Principal Site Reliability Engineer (SRE)

Senior • Remote

35,000 - 40,000 PLN/yr

Kraków, MA, Poland

Quick Facts

  • Role: Principal Site Reliability Engineer (SRE)

  • Focus: Hands-on reliability engineering for a business-critical product

Description

Design, implement, and operate highly available and scalable systems primarily on Kubernetes (AWS EKS), with a strong emphasis on operational excellence. Lead technical standards and guide engineers while applying GitOps and progressive delivery approaches (Terraform, ArgoCD, GitHub Actions; blue-green, canary releases, feature flagging). Participate in incident response, post-incident reviews, monitoring/observability improvements, and when required, production support.

Responsibilities

  • Design, operate, and troubleshoot Kubernetes clusters (AWS EKS), focusing on networking, scalability, security, and reliability

  • Architect and maintain fault-tolerant AWS infrastructure using Infrastructure as Code (Terraform)

  • Automate provisioning, deployment, and configuration processes using GitOps with ArgoCD and GitHub Actions

  • Define and enforce guardrails for infrastructure, applications, and databases to ensure secure and consistent operations

  • Implement and maintain monitoring and observability solutions with Prometheus, Grafana, and related tools

  • Build and evolve CI/CD pipelines and progressive delivery strategies

  • Collaborate with development teams to embed reliability and security best practices across the application lifecycle

  • Participate in incident response, post-incident reviews, continuous improvement, resilience testing, and chaos engineering

  • Design and manage secure networking solutions including AWS VPCs, Kubernetes networking, and firewalls

Requirements

  • 7+ years of commercial experience in SRE, systems engineering, infrastructure, or related roles

  • University degree in Computer Science or a related field

  • Strong hands-on Kubernetes experience (AWS EKS or similar), including networking, scaling, and security

  • Advanced knowledge of AWS services: EKS, EC2, CloudWatch, Route53, Aurora, S3

  • Proven experience with Terraform, ArgoCD, and GitHub Actions

  • Strong monitoring/observability and incident management experience (Prometheus, Grafana)

  • Strong scripting and automation skills in Python, Go, or Bash

  • Willingness to actively participate in production support when required

Benefits

  • Hands-on impact on system reliability, scalability, and technical direction

  • Opportunity to define operational standards, guardrails, and deployment strategy using modern GitOps and progressive delivery

  • Involvement in incident response, resilience testing, and chaos engineering to continuously improve production reliability

Similar jobs you might like