Unlock Full Resume Report

New offer - be the first one to apply!

September 16, 2026

Cloud Infrastructure / Site Reliability Engineer

Senior

132,000 - 227,400 USD/yr

Morrisville, NC

Quick Facts

  • Role: Cloud Infrastructure / Site Reliability Engineer (SRE)

  • Focus: Reliability, automation, monitoring, and incident response for SaaS/IaaS services built on microservices and Kubernetes

  • Support: Rotation-based on-call

Description

Operate at the intersection of development and operations, enhancing the cloud services lifecycle from design through deployment, operation, and continuous refinement. Maintain SaaS/IaaS environments by measuring and monitoring availability, latency, and overall system health, and by building automation for efficient cloud operations management. Support customer-centric cloud services with a focus on availability, performance, and security, and collaborate with cloud service provider teams including Azure.

Responsibilities

  • Identify opportunities for automation to reduce risk and improve time efficiency; build software for deployment automation, packaging, and monitoring visibility
  • Collaborate with Cloud Infrastructure Engineers and developers to maximize performance, reliability, and automation
  • Consult and influence developers on feature development and software architecture for scalability
  • Debug and troubleshoot bottlenecks across the software stack; provide advanced tier 2 and tier 3 support
  • Monitor, analyze, and measure system health, availability, and latency using Prometheus, Stackdriver, Elasticsearch, Grafana, and SolarWinds
  • Maintain and monitor deployment/orchestration of servers, Docker containers, databases, and backend infrastructure
  • Conduct root cause analysis (RCA) for complex production incidents involving OS, Networking, and Database; apply SRE best practices
  • Create runbooks and document system knowledge
  • Stay current on security protocols and resolve complex security issues
  • Use Atlassian tools and first-party cloud management tools to track and resolve issues by priority
  • Measure and monitor availability, latency, and system health to influence solution implementation decisions

Requirements

  • US citizenship (required due to potential need for Security Clearance)
  • 5+ years of scripting and infrastructure automation using PowerShell, Python, Go, or Ruby
  • Deep knowledge of containers, Kubernetes, serverless computing, and distributed systems design patterns
  • Knowledge of DevOps/SRE development methodologies
  • Proficiency in Linux/Unix and CoreOS
  • Experience with AWS, Azure, or Google Cloud
  • Ability to lead a scrum team; influence stakeholders to manage a product backlog and run sprints
  • On-call rotations and willingness to work odd hours

Benefits

  • Target salary range: $147,900–$220,000 USD (final compensation depends on location, qualifications, experience, and education)
  • Comprehensive benefits package: health insurance, life insurance, retirement/pension plans, paid time off, and additional leave options
  • Performance-based incentives and employee stock purchase plan / RSUs (region-dependent)
  • Hybrid working environment with some in-office and/or in-person expectations

Similar jobs you might like