Unlock Full Resume Report

New offer - be the first one to apply!

September 14, 2026

Cloud Operations Engineer - AI Cloud Ops

Mid

100,000 - 100,000 USD/yr

San Ramon, CA

Quick Facts

  • Role: Cloud Ops Engineer, AIOps (Cloud)
  • Work style: Hybrid (San Ramon, CA)

Description

Support the operation and maintenance of AWS-based cloud infrastructure. Deploy, manage, and operate applications on Kubernetes clusters (EKS or similar), including troubleshooting, production monitoring/alerting, and incident response. Contribute to infrastructure automation, improve CI/CD workflows for container-based releases, maintain reliability/performance, and leverage AI/ML tools and copilots to improve cloud operations efficiency and effectiveness.

Responsibilities

  • Operate and maintain cloud infrastructure across AWS environments
  • Deploy, manage, and operate applications on Kubernetes clusters (EKS or similar)
  • Troubleshoot Kubernetes issues (pod failures, scaling issues, networking, cluster health)
  • Assist in managing production systems (monitoring, alerting, incident response)
  • Troubleshoot and resolve infrastructure and application issues in a timely manner
  • Work with engineering teams to support deployments and improve CI/CD workflows, including container-based releases
  • Contribute to infrastructure automation using Terraform and Ansible
  • Maintain system reliability, availability, and performance through day-to-day operations
  • Participate in on-call rotation and incident response processes
  • Assist with capacity planning and cost optimization efforts
  • Partner with security teams to support compliance and security best practices
  • Contribute to improving observability via logs, metrics, and dashboards

AI-Driven Operations

  • Leverage AI/ML tools and copilots to improve cloud operations efficiency and effectiveness
  • Use AI-assisted solutions for incident triage, root cause analysis, alert noise reduction, and automation of repetitive tasks
  • Identify opportunities to apply AI to reduce operational toil and improve MTTR
  • Integrate AI-driven insights into monitoring, logging, and operational workflows

Requirements

  • 1–5 years of experience in cloud operations, DevOps, or infrastructure engineering
  • Ability to commute to the San Ramon office 2–3 times a week
  • Basic to strong understanding of AWS cloud services
  • Experience working in Linux/Unix environments
  • Familiarity with container technologies (Docker and/or Kubernetes)
  • Exposure to infrastructure-as-code tools (Terraform, Ansible, or similar)
  • Understanding of CI/CD concepts and tools
  • Familiarity with monitoring and logging tools (CloudWatch, ELK, Prometheus, Grafana)
  • Strong problem-solving skills and willingness to learn
  • Ability to work collaboratively in a team environment
  • Good communication skills and attention to detail
  • Hands-on experience using AI-powered tools (e.g., GitHub Copilot, AWS Kiro, Claude Code, Codex)
  • Understanding of AI applications in cloud operations (anomaly detection, automation, incident analysis)
  • Ability to use AI tools for troubleshooting, log analysis, and workflow improvement
  • Demonstrated curiosity and willingness to adopt AI-driven approaches

Nice To Have Qualifications

  • AWS certifications (Associate level preferred)
  • Exposure to scripting/programming (Python or Shell)
  • Experience in SaaS or cloud-based production environments
  • Exposure to AIOps or observability platforms with built-in AI capabilities

Benefits

  • Competitive health and wellness benefits (Medical, Dental, Vision, Life Insurance)
  • Flexible Spending Accounts
  • Flexible Time Off
  • 401(k) with employer match
  • Commuter Benefits
  • Opportunities for learning, mentorship, and career growth
  • Compensation (CIP): anticipated annual salary range of $100,000–$150,000; performance-based bonus and equity programs may apply

Similar jobs you might like