Unlock Full Resume Report

This position is no longer accepting applications

Positions open for more than 30 days are automatically closed and marked as expired

Don't let one closed door slow you down, here's your next move:

July 14, 2026

Member Of Technical Staff - Cloud Infrastructure

Senior • On-site

180,000 - 440,004 USD/yr

Palo Alto, CA

SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. The team is highly motivated, focused on engineering excellence, and operates with a flat organizational structure. Employees are expected to contribute directly, communicate effectively, and demonstrate strong prioritization and initiative.

About the Role

We are seeking a highly skilled Senior Infrastructure Engineer to join the US Government Team, focused on designing, building, and operating secure, scalable infrastructure for critical government projects. In this role, you will develop and manage training and inference clusters, as well as highly reliable applications, across bare metal, classified cloud, and hybrid cloud architectures. You will leverage expertise in Kubernetes and GPU hardware to deliver robust, secure systems that support large-scale AI workloads while meeting stringent federal compliance requirements.

This role demands a passion for automation, observability, and ensuring system integrity in a fast-paced, high-security environment.

Responsibilities

  • Develop and optimize software to provision and manage infrastructure across on-premise, virtual machine, and classified cloud environments.
  • Enhance the reliability, performance, and cost-effectiveness of infrastructure supporting large-scale AI and application workloads.
  • Collaborate with engineers to understand workload requirements and design compliant solutions for government projects.
  • Implement observability, monitoring, and security practices to ensure system integrity, availability, and confidentiality.
  • Manage storage infrastructure using Infrastructure-as-Code tools such as Pulumi, Terraform, or Ansible.
  • Drive reliability through incident management, postmortems, and the definition of SLAs and SLOs.
  • Work on-site in Palo Alto, CA or Washington, DC, with up to 50% travel required.

Basic Qualifications

  • Active Top Secret (TS) security clearance.
  • 5+ years of experience as an Infrastructure Engineer, Site Reliability Engineer, or similar role.
  • Experience building and maintaining reliable, scalable systems in secure or government environments.
  • Proficiency with Pulumi, Terraform, or Ansible.
  • Deep understanding of Kubernetes, including CNI, CRI, CSI, and related components.
  • Experience improving reliability through incident management, postmortems, and SLAs/SLOs.
  • Strong communication and documentation skills.

Preferred Skills and Experience

  • Experience installing and maintaining GPU hardware and drivers.
  • Experience optimizing Kubernetes for high-traffic deployments in classified or federal settings.
  • Familiarity with chaos engineering and capacity planning.
  • Proficiency with Kyverno, ArgoCD, or Go for infrastructure automation.
  • Security certifications such as CISSP or experience in secure federal environments.

Compensation and Benefits

  • $180,000 - $440,000 USD salary range.
  • Equity compensation.
  • Medical, vision, and dental coverage.
  • 401(k) retirement plan.
  • Short- and long-term disability insurance.
  • Life insurance and additional employee perks.

SpaceXAI is an equal opportunity employer. For details on data processing, view the Recruitment Privacy Notice.