Unlock Full Resume Report

New offer - be the first one to apply!

September 18, 2026

Senior Site Reliability Engineer

Senior

150,000 - 150,000 USD/yr

New York, NY

Quick Facts

  • Role: Senior Site Reliability Engineer

Description

Join an SRE team to design, build, operate, and evolve infrastructure powering internet-scale traffic, including cloud, Kubernetes clusters, and CI/CD platforms.

You will partner with development teams to improve reliability, scalability, and operational efficiency, ensuring safe and efficient deployments.

You’ll also handle production incidents, troubleshoot complex issues across the infrastructure stack, and drive continuous improvements based on lessons learned.

Role Description / Responsibilities

  • Operate and evolve cloud and CDN infrastructure with a focus on reliability, performance, scalability, and cost efficiency

  • Improve and optimize the CI/CD platform to enhance developer experience and deployment reliability

  • Drive the implementation and operation of multi-region Kubernetes clusters

  • Troubleshoot complex production and infrastructure issues with engineering teams using a holistic view across platforms

  • Stay current with emerging technologies (including agentic development and AI-assisted engineering) and identify opportunities to integrate them into platforms and engineering workflows

Requirements

  • Expert-level Infrastructure as a Service (IaaS) operations experience and Infrastructure as Code (IaC), especially Terraform and Atlantis

  • Extensive containerization and orchestration experience with Docker and Kubernetes

  • Strong AWS infrastructure operations experience, especially EKS, S3, EC2, and VPC

  • Ability to design and implement cost-effective solutions and proactively identify cost-optimization opportunities

  • Strong systems engineering fundamentals, with deep Linux and networking knowledge for diagnosing issues across OS, network, container, and cloud layers

  • Experience with CI/CD tools such as Jenkins or GitHub Actions

  • Experience with deployment tooling such as Helm, Spinnaker, or ArgoCD

  • Proficiency in one or more programming languages: Python, Java, Go, or Rust

  • Experience with observability and monitoring platforms such as Datadog and AWS CloudWatch, including production troubleshooting

  • Experience with agentic engineering and AI-assisted development practices

  • Strong problem-solving and communication skills

  • High autonomy and initiative for optimization, cost reduction, patching, upgrades, and managing technical debt

Benefits

  • Competitive compensation and a generous benefits package including health, wellness, and financial benefits

  • Pay range (New York): $170,000–$185,000 per year

Similar jobs you might like