July 14, 2026

Member Of Technical Staff - Cloud Infrastructure

Senior • On-site

180,000 - 440,004 USD/yr

Palo Alto, CA

SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. The team is highly motivated, focused on engineering excellence, and operates with a flat organizational structure. Employees are expected to contribute directly, communicate effectively, and demonstrate strong prioritization and initiative.

About the Role

We are seeking a highly skilled Senior Infrastructure Engineer to join the US Government Team, focused on designing, building, and operating secure, scalable infrastructure for critical government projects. In this role, you will develop and manage training and inference clusters, as well as highly reliable applications, across bare metal, classified cloud, and hybrid cloud architectures. You will leverage expertise in Kubernetes and GPU hardware to deliver robust, secure systems that support large-scale AI workloads while meeting stringent federal compliance requirements.

This role demands a passion for automation, observability, and ensuring system integrity in a fast-paced, high-security environment.

Responsibilities

  • Develop and optimize software to provision and manage infrastructure across on-premise, virtual machine, and classified cloud environments.
  • Enhance the reliability, performance, and cost-effectiveness of infrastructure supporting large-scale AI and application workloads.
  • Collaborate with engineers to understand workload requirements and design compliant solutions for government projects.
  • Implement observability, monitoring, and security practices to ensure system integrity, availability, and confidentiality.
  • Manage storage infrastructure using Infrastructure-as-Code tools such as Pulumi, Terraform, or Ansible.
  • Drive reliability through incident management, postmortems, and the definition of SLAs and SLOs.
  • Work on-site in Palo Alto, CA or Washington, DC, with up to 50% travel required.

Basic Qualifications

  • Active Top Secret (TS) security clearance.
  • 5+ years of experience as an Infrastructure Engineer, Site Reliability Engineer, or similar role.
  • Experience building and maintaining reliable, scalable systems in secure or government environments.
  • Proficiency with Pulumi, Terraform, or Ansible.
  • Deep understanding of Kubernetes, including CNI, CRI, CSI, and related components.
  • Experience improving reliability through incident management, postmortems, and SLAs/SLOs.
  • Strong communication and documentation skills.

Preferred Skills and Experience

  • Experience installing and maintaining GPU hardware and drivers.
  • Experience optimizing Kubernetes for high-traffic deployments in classified or federal settings.
  • Familiarity with chaos engineering and capacity planning.
  • Proficiency with Kyverno, ArgoCD, or Go for infrastructure automation.
  • Security certifications such as CISSP or experience in secure federal environments.

Compensation and Benefits

  • $180,000 - $440,000 USD salary range.
  • Equity compensation.
  • Medical, vision, and dental coverage.
  • 401(k) retirement plan.
  • Short- and long-term disability insurance.
  • Life insurance and additional employee perks.

SpaceXAI is an equal opportunity employer. For details on data processing, view the Recruitment Privacy Notice.

Similar jobs you might like

Technology

xAI

Software Engineer - Platform Security

Mid

On-site

Palo Alto, CA

99,996 - 258,000 USD/yr

🏢 Summary: Software Engineer role focused on platform security, building AI-driven security tools and securing Kubernetes-based infrastructure and applications. The position involves designing scalable backend systems, identifying vulnerabilities, and driving secure engineering practices in a fast-paced environment. 🗂️ Requirements: 3+ years of experience in fast-paced technology environments, Expertise in Python, Rust, or Go, Experience building scalable tools or systems from scratch, Proficiency in scalable backend architecture design, Familiarity with security testing frameworks, Experience with Docker and Kubernetes, Knowledge of SBOM management and dependency scanning, Strong problem-solving and clean coding skills 📃 Skills: Python, Rust, Go, Kubernetes, Docker, Grok, BurpSuite, OWASP, SAST, DAST, SBOM 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: We are seeking a talented and driven Software Engineer to join Platform Security team, where you will build cutting-edge security solutions to protect our Kubernetes-based infrastructure and advance secure AI-driven systems. In this role, you will design and implement AI-powered security tools, proactively address vulnerabilities, and champion secure engineering practices across the organization. Ideal candidates are passionate about impactful innovation, excel at writing clean, efficient code, and thrive in fast-paced environments to support SpaceXAI's mission of creating a trusted and secure global digital platform. RESPONSIBILITIES: - Design and build AI-driven security tooling and agents using Grok to identify, analyze, and mitigate vulnerabilities in the platform infrastructure and customer facing application(s) - Proactively identify security problems to solve and own the design and implementation end-to-end - Collaborate and be a security champion while driving technical decisions across the organization BASIC QUALIFICATIONS: - 3+ years of experience in fast-paced, high-impact environments, ideally at startups or tech-driven companies. - Expertise in Python, Rust, or Go, with strong problem-solving skills and a focus on clean, efficient code. - Certifications like CISA, CRISC, CGEIT, Security+, CASP+, or similar preferred. - Proven experience building tools or systems from scratch, with a focus on scalable solutions. - Proficiency in designing scalable backend architectures to support secure systems. - Familiarity with security testing frameworks (e.g., Burp Suite, OWASP ZAP, SAST/DAST). - Experience with Docker and Kubernetes for deploying and securing containerized applications. - Knowledge of software supply chain tools, including SBOM management and dependency scanning. PREFERRED SKILLS AND EXPERIENCE: - Experience developing AI-driven security tools or integrating AI into security workflows. - Familiarity with Kubernetes-based environments and securing cloud-native infrastructure. - Proven ability to drive technical decisions and influence security practices across teams. - A passion for challenging the status quo and building transformative security solutions. - Strong collaboration skills, with experience working in dynamic, cross-functional teams. - A sense of humor and adaptability to thrive in a fast-paced, mission-driven environment. COMPENSATION AND BENEFITS: - $100,000 - $258,000 USD - Base salary is just one part of the total rewards package, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks. SpaceXAI is an equal opportunity employer. For details on data processing, view the Recruitment Privacy Notice.

Technology

SpaceXAI

Member of Technical Staff - Recommendation Systems

Senior

On-site

Palo Alto, CA

180,000 - 440,004 USD/yr

🏢 Summary: Applied Engineer role focused on building and scaling recommendation systems, ranking algorithms, and AI-driven search technologies for products serving hundreds of millions of users. The position involves developing machine learning infrastructure, real-time experimentation, and large-scale deep learning applications using modern AI frameworks and data systems. 🗂️ Requirements: Knowledge of Kafka, Clickhouse, and Spark, Experience implementing recommender systems at industrial scale, Experience with deep learning applications at industrial scale, Proficiency in JAX or PyTorch, Ability to design scalable machine learning systems, Strong communication skills, Hands-on engineering approach 📃 Skills: Kafka, Clickhouse, Spark, JAX, PyTorch, CUDA, Python, AI, ML 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: We're seeking exceptional Applied engineers to join a high-priority project that approximately 600 million monthly users use. This is an exciting opportunity for individuals with an engineer or scientist background to apply their skills to recommendation systems, ranking algorithms, search technologies, and many other systems. You'll work at the intersection of advanced AI development and real-world impact, enhancing the ability to connect users with relevant content, accounts, and experiences. RESPONSIBILITIES: - Designing and architecting recommendation algorithms across various product surfaces - Leverage all of xAI's infra and AI stacks to dramatically enhance the user experience - Write data pipelines and training jobs that continuously learn from product data - Iterate and improve the algorithm by gathering user feedback in real time through experimentation - Ensuring scalability and efficiency of machine learning systems BASIC QUALIFICATIONS: - Knowledge of data infrastructure like Kafka, Clickhouse, and Spark - Experienced in implementing recommender systems and/or deep learning applications at industrial scale - Skilled in one or more DL software frameworks such as JAX or PyTorch - Exceptional candidates may be experienced in writing CUDA kernels COMPENSATION AND BENEFITS: - $180,000 - $440,000 USD - Base salary is one part of the total rewards package and includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short and long-term disability insurance, life insurance, and additional discounts and perks. SpaceXAI is an equal opportunity employer. For details on data processing, view the Recruitment Privacy Notice.

Technology

SpaceXAI

Senior IAM Engineer

Senior

On-site

Washington, DC

99,996 - 258,000 USD/yr

🏢 Summary: Senior IAM Engineer role focused on designing, implementing, and optimizing enterprise identity and access management solutions across cloud, on-premises, and SaaS environments. The position involves leading federation, IAM automation, access governance, and cloud IAM initiatives while collaborating with security and engineering teams. The role requires deep expertise in IAM platforms, protocols, scripting, and cloud technologies in a fast-paced environment. 🗂️ Requirements: 6+ years of IAM experience, Expertise in SAML, Expertise in OIDC, Expertise in OAuth2, Expertise in SCIM, Experience with SailPoint, Experience with Okta, Experience with PingOne, Experience with Microsoft Entra ID, Knowledge of Zero Trust, Knowledge of Least Privilege, Knowledge of ABAC, Knowledge of ITDR, Experience with AWS, Experience with GCP, Experience with Terraform, Experience with CI/CD, Programming skills in Python, Programming skills in JavaScript, Programming skills in TypeScript, Experience designing federation solutions, Ability to lead IAM initiatives end-to-end 📃 Skills: IAM, SAML, OIDC, OAuth2, SCIM, SailPoint, Okta, PingOne, Entra, RBAC, PAM, Terraform, Python, JavaScript, TypeScript, AWS, GCP, CI/CD, ABAC, ITDR 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: We are looking for a Senior Identity and Access Management (IAM) Engineer to play a critical role in designing, implementing, and maturing our enterprise IAM ecosystem. As a Senior IAM Engineer, you will serve as a key subject matter expert (SME) within the broader Information Security organization, driving secure, scalable, and user-friendly identity solutions in a fast-paced, high-growth environment. You will own complex IAM initiatives end-to-end — from strategy and architecture to implementation and optimization — while collaborating closely with Security, Infrastructure, Application, and Engineering teams. This is a high-impact, hands-on role suited for someone who thrives in ambiguity and takes extreme ownership of outcomes. RESPONSIBILITIES: - Design and implement enterprise-grade IAM solutions across on-prem, cloud, and SaaS environments. - Architect and deliver end-to-end federation strategies using modern protocols (SAML, OIDC, OAuth 2.0, SCIM). - Lead integration and optimization of key IAM platforms including SailPoint, Okta, PingOne, Microsoft Entra ID, and other tools. - Drive Identity Lifecycle Management (LCM), Role-Based Access Control (RBAC), Privileged Access Management (PAM), and Just-In-Time access initiatives. - Develop automation scripts and Infrastructure as Code (Terraform) to improve provisioning, governance, and operational efficiency. - Serve as the primary IAM SME, providing technical guidance, troubleshooting complex issues, and mentoring other team members. - Work closely with InfoSec, Compliance, and Engineering teams to meet audit, regulatory, and security requirements. - Continuously improve IAM architecture, processes, and capabilities in a dynamic, fast-moving organization. - Support cloud IAM strategies across AWS, GCP, and other platforms. - Participate in on-call rotation and incident response related to identity systems. BASIC QUALIFICATIONS: - 6+ years of hands-on experience in Identity and Access Management. - Thorough understanding of IAM principles and modern security concepts, including Zero Trust, Least Privilege, Adaptive/Risk-Based Authentication, Attribute-Based Access Control (ABAC), Continuous Access Evaluation, and Identity Threat Detection and Response (ITDR). - Strong expertise in IAM protocols and standards: SAML, OIDC, OAuth 2.0, SCIM, and other federation protocols. - Strong expertise in IAM and federation protocols, with proven experience in architecting and implementing end-to-end federation solutions. - Hands-on experience with major IAM platforms, particularly SailPoint, Okta, PingOne, and Microsoft identity solutions. - Strong working knowledge of cloud platforms (AWS and GCP), Cloud IAM services, Terraform, and CI/CD practices. - Strong scripting and programming skills in Python and JavaScript/TypeScript, with extensive experience developing IAM workflows, automation, and integrations. - Demonstrated ability to lead technical initiatives and drive projects from design through implementation. This high-visibility role requires the ability to thrive in a fast-paced and ambiguous environment. The ideal candidate demonstrates extreme ownership, a collaborative no-ego mindset, and is comfortable taking the lead. This is a significant opportunity to shape the future of identity and security at the company. COMPENSATION AND BENEFITS $100,000 - $258,000 USD Base salary is just one part of our total rewards package, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks. SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Technology

SpaceXAI

Senior IAM Engineer

Senior

On-site

Austin, TX

99,996 - 258,000 USD/yr

🏢 Summary: Senior IAM Engineer role focused on designing, implementing, and optimizing enterprise identity and access management solutions across cloud, on-prem, and SaaS environments. The position involves leading federation, IAM automation, access control, and cloud identity initiatives while collaborating with security and engineering teams. This is a hands-on, high-impact role in a fast-paced environment requiring deep IAM expertise and strong technical leadership. 🗂️ Requirements: 6+ years of IAM experience, Expertise in SAML, Expertise in OIDC, Expertise in OAuth2, Expertise in SCIM, Experience with SailPoint, Experience with Okta, Experience with PingOne, Experience with Microsoft Entra ID, Knowledge of Zero Trust, Knowledge of Least Privilege, Knowledge of ABAC, Knowledge of ITDR, Experience with AWS, Experience with GCP, Experience with Terraform, Experience with CI/CD, Programming skills in Python, Programming skills in JavaScript, Programming skills in TypeScript, Experience designing federation solutions, Experience with IAM automation and integrations, Ability to lead technical initiatives 📃 Skills: IAM, SAML, OIDC, OAuth2, SCIM, SailPoint, Okta, PingOne, Entra, RBAC, PAM, ABAC, ITDR, Terraform, Python, JavaScript, TypeScript, AWS, GCP, CI/CD 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: We are looking for a Senior Identity and Access Management (IAM) Engineer to play a critical role in designing, implementing, and maturing our enterprise IAM ecosystem. As a Senior IAM Engineer, you will serve as a key subject matter expert (SME) within the broader Information Security organization, driving secure, scalable, and user-friendly identity solutions in a fast-paced, high-growth environment. You will own complex IAM initiatives end-to-end — from strategy and architecture to implementation and optimization — while collaborating closely with Security, Infrastructure, Application, and Engineering teams. This is a high-impact, hands-on role suited for someone who thrives in ambiguity and takes extreme ownership of outcomes. RESPONSIBILITIES: - Design and implement enterprise-grade IAM solutions across on-prem, cloud, and SaaS environments. - Architect and deliver end-to-end federation strategies using modern protocols (SAML, OIDC, OAuth 2.0, SCIM). - Lead integration and optimization of key IAM platforms including SailPoint, Okta, PingOne, Microsoft Entra ID, and other tools. - Drive Identity Lifecycle Management (LCM), Role-Based Access Control (RBAC), Privileged Access Management (PAM), and Just-In-Time access initiatives. - Develop automation scripts and Infrastructure as Code (Terraform) to improve provisioning, governance, and operational efficiency. - Serve as the primary IAM SME, providing technical guidance, troubleshooting complex issues, and mentoring other team members. - Work closely with InfoSec, Compliance, and Engineering teams to meet audit, regulatory, and security requirements. - Continuously improve IAM architecture, processes, and capabilities in a dynamic, fast-moving organization. - Support cloud IAM strategies across AWS, GCP, and other platforms. - Participate in on-call rotation and incident response related to identity systems. BASIC QUALIFICATIONS: - 6+ years of hands-on experience in Identity and Access Management. - Thorough understanding of IAM principles and modern security concepts, including Zero Trust, Least Privilege, Adaptive/Risk-Based Authentication, Attribute-Based Access Control (ABAC), Continuous Access Evaluation, and Identity Threat Detection and Response (ITDR). - Strong expertise in IAM protocols and standards: SAML, OIDC, OAuth 2.0, SCIM, and other federation protocols. - Strong expertise in IAM and federation protocols, with proven experience in architecting and implementing end-to-end federation solutions. - Hands-on experience with major IAM platforms, particularly SailPoint, Okta, PingOne, and Microsoft identity solutions. - Strong working knowledge of cloud platforms (AWS and GCP), Cloud IAM services, Terraform, and CI/CD practices. - Strong scripting and programming skills in Python and JavaScript/TypeScript, with extensive experience developing IAM workflows, automation, and integrations. - Demonstrated ability to lead technical initiatives and drive projects from design through implementation. This high-visibility role requires the ability to thrive in a fast-paced and ambiguous environment. The ideal candidate demonstrates extreme ownership, a collaborative no-ego mindset, and is comfortable taking the lead. This is a significant opportunity to shape the future of identity and security at the company. COMPENSATION AND BENEFITS $100,000 - $258,000 USD Base salary is just one part of our total rewards package, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks. SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Technology

SpaceXAI

Senior IAM Engineer

Senior

On-site

Palo Alto, CA

99,996 - 258,000 USD/yr

🏢 Summary: Senior IAM Engineer role focused on designing, implementing, and optimizing enterprise identity and access management solutions across cloud, on-premises, and SaaS environments. The position involves leading federation, IAM automation, RBAC, PAM, and cloud IAM initiatives while collaborating with security and engineering teams in a fast-paced environment. Candidates will work hands-on with modern IAM platforms, protocols, and Infrastructure as Code technologies. 🗂️ Requirements: 6+ years of IAM experience, Expertise in SAML, Expertise in OIDC, Expertise in OAuth 2.0, Expertise in SCIM, Experience with SailPoint, Experience with Okta, Experience with PingOne, Experience with Microsoft Entra ID, Knowledge of Zero Trust, Knowledge of Least Privilege, Knowledge of ABAC, Knowledge of ITDR, Experience with AWS, Experience with GCP, Experience with Terraform, Experience with CI/CD, Programming skills in Python, Programming skills in JavaScript, Programming skills in TypeScript, Experience designing federation solutions, Ability to lead technical initiatives 📃 Skills: IAM, SAML, OIDC, OAuth, SCIM, SailPoint, Okta, PingOne, Entra, RBAC, PAM, Terraform, AWS, GCP, Python, JavaScript, TypeScript, CI/CD, ABAC, ITDR 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: We are looking for a Senior Identity and Access Management (IAM) Engineer to play a critical role in designing, implementing, and maturing our enterprise IAM ecosystem. As a Senior IAM Engineer, you will serve as a key subject matter expert (SME) within the broader Information Security organization, driving secure, scalable, and user-friendly identity solutions in a fast-paced, high-growth environment. You will own complex IAM initiatives end-to-end — from strategy and architecture to implementation and optimization — while collaborating closely with Security, Infrastructure, Application, and Engineering teams. This is a high-impact, hands-on role suited for someone who thrives in ambiguity and takes extreme ownership of outcomes. RESPONSIBILITIES: - Design and implement enterprise-grade IAM solutions across on-prem, cloud, and SaaS environments. - Architect and deliver end-to-end federation strategies using modern protocols (SAML, OIDC, OAuth 2.0, SCIM). - Lead integration and optimization of key IAM platforms including SailPoint, Okta, PingOne, Microsoft Entra ID, and other tools. - Drive Identity Lifecycle Management (LCM), Role-Based Access Control (RBAC), Privileged Access Management (PAM), and Just-In-Time access initiatives. - Develop automation scripts and Infrastructure as Code (Terraform) to improve provisioning, governance, and operational efficiency. - Serve as the primary IAM SME, providing technical guidance, troubleshooting complex issues, and mentoring other team members. - Work closely with InfoSec, Compliance, and Engineering teams to meet audit, regulatory, and security requirements. - Continuously improve IAM architecture, processes, and capabilities in a dynamic, fast-moving organization. - Support cloud IAM strategies across AWS, GCP, and other platforms. - Participate in on-call rotation and incident response related to identity systems. BASIC QUALIFICATIONS: - 6+ years of hands-on experience in Identity and Access Management. - Thorough understanding of IAM principles and modern security concepts, including Zero Trust, Least Privilege, Adaptive/Risk-Based Authentication, Attribute-Based Access Control (ABAC), Continuous Access Evaluation, and Identity Threat Detection and Response (ITDR). - Strong expertise in IAM protocols and standards: SAML, OIDC, OAuth 2.0, SCIM, and other federation protocols. - Strong expertise in IAM and federation protocols, with proven experience in architecting and implementing end-to-end federation solutions. - Hands-on experience with major IAM platforms, particularly SailPoint, Okta, PingOne, and Microsoft identity solutions. - Strong working knowledge of cloud platforms (AWS and GCP), Cloud IAM services, Terraform, and CI/CD practices. - Strong scripting and programming skills in Python and JavaScript/TypeScript, with extensive experience developing IAM workflows, automation, and integrations. - Demonstrated ability to lead technical initiatives and drive projects from design through implementation. This high-visibility role requires the ability to thrive in a fast-paced and ambiguous environment. The ideal candidate demonstrates extreme ownership, a collaborative no-ego mindset, and is comfortable taking the lead. This is a significant opportunity to shape the future of identity and security at the company. COMPENSATION AND BENEFITS $100,000 - $258,000 USD Base salary is just one part of the total rewards package, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks. SpaceXAI is an equal opportunity employer. For details on data processing, view the Recruitment Privacy Notice.

Technology

SpaceXAI

Network Engineer - ML Infrastructure (High-Speed Interconnects)

Senior

On-site

Palo Alto, CA

180,000 - 440,004 USD/yr

🏢 Summary: Senior ML Infrastructure Engineer role focused on designing and optimizing high-speed copper and optical interconnects for large-scale AI training and inference clusters. The position involves end-to-end ownership of network fabric architecture, hardware validation, diagnostics, automation, and collaboration with vendors and ML teams. Candidates will work on cutting-edge AI networking technologies including photonics, SerDes, and next-generation optical systems. 🗂️ Requirements: 8+ years of experience with high-speed copper and optical interconnects, Master’s or PhD in Electrical Engineering, Photonics, or Physics, Expert knowledge of PAM4 SerDes, equalization, jitter, and crosstalk, Operational understanding of FEC, Retimers, TIAs, and Drivers, Experience with optical link budget analysis and diagnostics, Expertise in transceiver components and failure characterization, Knowledge of thermal, mechanical, power, and signal integrity constraints, Knowledge of SiPh design process, yield improvement, and reliability testing, Familiarity with CPO technologies and associated risks, Familiarity with supply chains, ODMs, and contract manufacturers, Strong problem-solving skills in fast-paced environments 📃 Skills: PAM4, SerDes, FEC, Retimers, TIAs, Drivers, SiPh, CPO, Photonics, IEEE, CMIS, DSP, VCSEL, microLED, THz, TDECQ, OMA, Python, Automation, Telemetry 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: xAI is building at a furious pace with the latest compute and switching hardware to help people understand the universe. We are looking for exceptional ML Infrastructure Engineers with deep expertise in high-speed interconnect technologies to design, build, and optimize the network fabric that powers large-scale AI training and inference clusters. This strategic role will drive innovation in high-bandwidth, low-latency, power-efficient interconnects critical for AI/ML clusters based on advanced computing platforms. You will have the opportunity to work on all modalities of interconnects connecting GPUs and switches both inside and between data centers, including our primary front and backend networks that train Grok and that customers use for inference. Engineers will own all aspects from design and development to build and operations. You will be expected to define and improve team processes and to contribute to scaling and maintenance efforts. You will focus on the physical layer and system-level integration of copper (ACC, AEC, CPC) and optical (FRO, LRO/TRO, LPO, AOC, CPO) interconnects that directly determine the performance, power efficiency, scale, and cost of next-generation AI/ML clusters. This is a highly technical, hands-on role bridging ML cluster requirements with cutting-edge interconnect hardware — ideal for engineers who love both large-scale AI systems and the physics/engineering of 200G+ SerDes, PAM4, photonics, signal integrity and diagnostics. RESPONSIBILITIES: - Design, validate, and productize high-speed copper and optical connectivity solutions for AI clusters (100k+ GPU scale). - Own vendor due diligence and onboarding for new 1.6T products including AEC and pluggable optical transceivers (DR4/8, FR4) including rigorous bring-up & characterization. - Investigate the opportunity for LPO and LRO in our network. - Evaluate early co-packaged and near-packaged engines for switches and GPUs. - Pathfinding for new interconnect modalities including VCSEL, microLED, THz radio-based solutions to improve network economics and reliability. - Work closely with vendors (transceiver, cable, SerDes, DSP, silicon photonics foundries) to influence roadmaps and ensure timely delivery of next-gen solutions. - Collaborate with ML training teams to translate workload communication patterns into concrete interconnect topology and optical reconfigurability requirements. - Perform system-level simulation of end-to-end fabric performance. - Drive failure analysis, root cause, and corrective actions for interconnect-related issues in production clusters through fleet-level metrics gathering and analysis. - Contribute to internal tooling and automation for interconnect health monitoring, telemetry, diagnostics, remediation and automated qualification pipelines. - Stay current with industry standards (OIF CMIS, IEEE) and emerging technologies (multi-core/hollow-core fiber, 448G SerDes, TFLN, ring resonators) BASIC QUALIFICATIONS: - At least 8+ years of hands-on experience in designing, deploying and operating high-speed copper and optical interconnects, preferably in a module design role or in a hyperscale datacenter environment. - Master's or PhD degree in Electrical Engineering, Photonics or Physics. - Deep knowledge of PAM4 SerDes performance, equalization, jitter, crosstalk. - Solid operational understanding of FEC, Retimers, TIAs and Drivers. - Deep knowledge of optical link budget analysis and performance metrics including TDECQ, OMA, Tcode, stressed receiver sensitivity and associated diagnostics. - Expertise in transceiver components including CW lasers, SiPh PICs, EML, DSP, passive subassemblies, their failure modes and characterization. - Knowledge of thermal, mechanical, power, signal integrity constraints in dense hardware. - Knowledge of SiPh design process, yield improvement and reliability testing. - Familiarity with CPO technologies and challenges/risk areas. - Familiarity with subcomponent supply chains and global manufacturers, ODMs and CMs. - Strong problem-solving skills and ability to thrive in a fast-paced, ambiguous setting. COMPENSATION AND BENEFITS: $180,000 - $440,000 USD Base salary is just one part of our total rewards package, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks. SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Technology

xAI

Site Reliability Engineer - Cybersecurity

Senior

On-site

Palo Alto, CA

🏢 Summary: Cybersecurity / SRE role focused on securing and maintaining the reliability of a large-scale fintech platform operating in hybrid cloud environments. The position emphasizes Kubernetes and container security, SIEM management, CI/CD protection, and automation using Python and infrastructure-as-code tools. Candidates will work on mission-critical distributed systems, ensuring regulatory compliance and resilient security operations at scale. 🗂️ Requirements: Experience securing hybrid AWS/on-premises environments, Strong proficiency in Python, Strong proficiency in Terraform, Strong proficiency in Puppet, Deep expertise in Kubernetes, Experience with container security, Hands-on experience with GitHub Actions, Experience with Prometheus, Experience with Grafana, Experience with CloudWatch, Experience with Karma, Experience managing and integrating Wazuh, Experience with security scanning tools (Semgrep, Trivy, Falco), Experience with IAM and security posture management, Ability to comply with PCI and NIST CSF standards, Located in SF Bay Area or willing to relocate 📃 Skills: AWS, IAM, Python, Terraform, Puppet, Kubernetes, Docker, GitHub, Prometheus, Grafana, CloudWatch, Karma, Wazuh, Semgrep, Trivy, Falco, PCI, NIST, CI/CD 🏢 Description: ABOUT xAI xAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: The Cybersecurity / SRE team is focused on ensuring the security and reliability of X Money. This role will primarily focus on the X Money platform but will also cross over with the X Social platform. The ideal candidate will have experience in the banking, money transmission, and P2P payments industry. We emphasize working with large distributed systems and security platforms at scale, with an automation-first mindset. You'll be responsible for securing and maintaining the reliability of X Money's infrastructure. You'll work closely with cross-functional teams to enhance security measures, improve system resilience, and implement best practices. RESPONSIBILITIES: - Build and secure mission-critical applications in a hybrid cloud environment. - Manage identities and roles effectively. - Monitor and remediate infrastructure to comply with regulations and best practices (e.g., PCI, NIST CSF). - Maintain a SIEM and all data pipelines needed for reliable alerting. - Design and implement secure container standards and automation to enable frictionless developer workflows. - Maintain Kubernetes security aligned with current best practices. - Build, deploy, and maintain security operations infrastructure using Python, Terraform, and Puppet. - Secure and enhance CI/CD pipelines. - Integrate and maintain code scanning platforms. - Develop dashboards and alerts from security metrics. - Own security projects: identify issues and implement solutions. - Apply critical analysis and problem-solving skills. BASIC QUALIFICATIONS: - Proven experience securing hybrid AWS/on-premises environments, including IAM and overall security posture. - Strong proficiency in Python, Terraform, and Puppet. - Certifications like CISA, CRISC, CGEIT, Security+, CASP+, or similar preferred. - Deep expertise in Kubernetes and container security. - Hands-on expertise building GitHub Actions and workflows. - Extensive experience with Prometheus, Grafana, CloudWatch, and Karma. - Well versed in management and integrations of Wazuh. - Hands-on experience with security scanning tools (Semgrep, Trivy, Falco). - Proactive mindset with strong ownership and problem-solving skills. - Excellent critical thinking and analytical abilities. - Located in the SF Bay Area or willing to relocate. COMPENSATION AND BENEFITS: $180,000 - $440,000 USD Base salary is just one part of our total rewards package, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks. xAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Technology

SpaceXAI

Member of Technical Staff - RL Training Framework

Senior

On-site

Palo Alto, CA

180,000 - 440,004 USD/yr

🏢 Summary: Engineering role focused on developing and optimizing reinforcement learning training infrastructure for large-scale workloads, from experimentation to production systems. The position involves improving scalability, observability, and end-to-end training performance in distributed environments. 🗂️ Requirements: Experience with large-scale distributed systems, Ability to debug and optimize system efficiency, Problem-solving across all levels of the stack, Proficiency in Python, Proficiency in Jax, Rust, or C++, Strong communication skills 📃 Skills: Python, Jax, Rust, C++, RL, LLM 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: The RL infrastructure team is looking for an engineer to help develop our RL training framework. RESPONSIBILITIES: - Design and implement the systems backing all RL workloads, from small scale ablations to production training runs - Profile, debug, and optimize end-to-end training performance - Improve scalability and observability of the RL stack BASIC QUALIFICATIONS: - Experience building, debugging, and optimizing efficiency of large-scale distributed systems - Comfortable diving into unfamiliar areas and solving problems at all levels of the stack - Proficiency in Python, Jax, Rust, and/or C++ PREFERRED SKILLS AND EXPERIENCE: - Experience with large scale LLM training infrastructure - Strong knowledge of reinforcement learning techniques - Experience with RL numerics COMPENSATION AND BENEFITS: - $180,000 - $440,000 USD - Equity package - Medical, vision, and dental coverage - 401(k) retirement plan - Short and long-term disability insurance - Life insurance - Additional discounts and perks SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Technology

xAI

Software Engineer - Training/Inference (C++)

Senior

On-site

Palo Alto, CA

🏢 Summary: High-impact inference engineering role focused on building and optimizing large-scale distributed model serving systems for Grok. The position involves low-level GPU and inference optimization, scalable infrastructure development, and ensuring high-performance, reliable AI serving at massive scale. Candidates will work across the full inference stack, from orchestration to GPU kernels and CI/CD systems. 🗂️ Requirements: Deep systems programming in C/C++ or Rust, Experience with large-scale high-concurrency production serving, Experience with GPU inference engines, Strong knowledge of batching, caching, load balancing, and parallelism, Experience with GPU kernel optimization and code generation, Knowledge of quantization, speculative decoding, distillation, and low-precision numerics, Experience with testing, benchmarking, and reliability engineering for inference services, Experience designing and implementing CI/CD infrastructure for inference 📃 Skills: C, C++, Rust, vLLM, SGLang, Triton, TensorRT-LLM, GPU, CI/CD, Quantization, Distillation, Parallelism, Caching, Benchmarking, Load-balancing, Autoscaling 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. About the Role: We are building the high-performance inference platform that serves Grok to millions of users every day with lightning speed and perfect reliability. As a Member of Technical Staff - Inference, you will design and optimize large-scale model serving systems end-to-end. You will own everything from distributed infrastructure (global KV cache, continuous batching, load balancing, auto-scaling) to deep low-level optimizations (GPU kernels, quantization, speculative decoding, tail latency). This is a high-impact role where your work directly determines how fast and reliably users interact with Grok at massive scale. Responsibilities: - Architect and implement scalable distributed infrastructure for model serving (load balancing, auto-scaling, batch scheduling, global KV cache). - Optimize latency and throughput of model inference under real production workloads. - Build reliable, high-concurrency serving systems that serve billions of users with 100% uptime, 0% error rate, and excellent tail latency. - Benchmark, fine-tune, and accelerate inference engines (including low-level GPU kernel work and code generation). - Develop custom tools to trace, replay, and fix issues across the full stack — from orchestration down to GPU kernels. - Create robust CI/CD infrastructure for seamless endpoint deployment, image publishing, and inference engine updates. - Accelerate research on scaling test-time compute, RL rollout, and model-hardware co-design for next-generation systems. Basic Qualifications: - Deep low-level systems programming (C/C++ or Rust) - Experience with large-scale, high-concurrent production serving. - Experience with GPU inference engines (vLLM, SGLang, Triton, TensorRT-LLM, etc.). - Strong background in system optimizations: batching, caching, load balancing, parallelism. - Low-level inference optimizations: GPU kernels, code generation. - Algorithmic inference optimizations: quantization, speculative decoding, distillation, low-precision numerics. - Experience with testing, benchmarking, and reliability of inference services. - Experience designing and implementing CI/CD infrastructure for inference. Compensation and Benefits: - $180,000 - $440,000 USD - Base salary is just one part of the total rewards package, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks. SpaceXAI is an equal opportunity employer. For details on data processing, view the Recruitment Privacy Notice.

Technology

SpaceXAI

Software Engineer - Kernels/CUDA (C++)

Senior

On-site

Palo Alto, CA

129,996 - 234,996 USD/yr

🏢 Summary: Role focused on building and optimizing large-scale GPU supercomputing infrastructure for AI training and inference. The position involves low-level GPU kernel optimization, distributed compute systems, and full-stack performance tuning across hardware and platform layers. Candidates will work closely with AI research teams to improve scalability, reliability, and training speed. 🗂️ Requirements: Deep low-level systems programming experience, Experience with large-scale GPU clusters, Hands-on GPU kernel optimization, Experience with distributed compute infrastructure, Ability to optimize memory-bound and compute-bound workloads, Strong communication skills 📃 Skills: CUDA, C, C++, PTX, SASS, CUTLASS, Nsight, Linux, GPU, Tensor 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: We are building one of the world's largest AI supercomputers from the ground up. As part of the Compute Infrastructure team, you will own both the raw GPU supercomputer and the platform layer that runs on top of it. You will work across the full stack — from low-level GPU kernel optimizations and Linux kernel internals to massive-scale orchestration and virtualization — to make training and inference as fast, reliable, and scalable as possible. This is a broad, high-impact role that combines hardcore supercompute and compute infrastructure work. Your contributions will directly accelerate Grok's training speed and overall AI progress. RESPONSIBILITIES: - Design, build, and optimize massive GPU clusters for extreme-scale training and inference workloads - Develop and tune low-level CUDA kernels (GeMM, Attention, etc.), using CUTLASS, Tensor Cores, and Nsight for maximum performance - Profile, debug, and eliminate bottlenecks across GPU memory hierarchy, networking fabric, filesystems, and multi-GPU operation - Collaborate closely with AI research teams to deliver production-grade performance and scalability PREFERRED SKILLS AND EXPERIENCE: - Deep low-level systems programming (C/C++/PTX/SASS) - Strong experience with large-scale GPU clusters or distributed compute infrastructure at production scale - Hands-on work with GPU kernel optimization (CUTLASS, custom kernels, Nsight profiling) - Track record of building or running high-performance infrastructure for AI workloads (training or inference platforms) - Ability to reason from first principles and optimize for both memory-bound and compute-bound scenarios COMPENSATION AND BENEFITS: - $180,000 - $440,000 USD - Equity - Medical, vision, and dental coverage - 401(k) retirement plan - Short and long-term disability insurance - Life insurance - Additional discounts and perks SpaceXAI is an equal opportunity employer. For details on data processing, view the Recruitment Privacy Notice.