July 14, 2026

Network Engineer - ML Infrastructure (High-Speed Interconnects)

Senior • On-site

180,000 - 440,004 USD/yr

Palo Alto, CA

SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity.

We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence.

About the Role

xAI is building at a furious pace with the latest compute and switching hardware to help people understand the universe. We are looking for exceptional ML Infrastructure Engineers with deep expertise in high-speed interconnect technologies to design, build, and optimize the network fabric that powers large-scale AI training and inference clusters.

You will work on all modalities of interconnects connecting GPUs and switches both inside and between data centers, including primary front and backend networks used for AI training and inference. Engineers own all aspects from design and development to build and operations.

This role focuses on the physical layer and system-level integration of copper and optical interconnects that directly determine the performance, power efficiency, scale, and cost of next-generation AI/ML clusters.

Responsibilities

  • Design, validate, and productize high-speed copper and optical connectivity solutions for AI clusters.
  • Own vendor due diligence and onboarding for new 1.6T products including AEC and pluggable optical transceivers.
  • Investigate LPO and LRO opportunities within the network.
  • Evaluate co-packaged and near-packaged engines for switches and GPUs.
  • Research new interconnect modalities including VCSEL, microLED, and THz radio-based solutions.
  • Collaborate with vendors to influence product roadmaps and ensure delivery of next-generation solutions.
  • Work with ML training teams to define interconnect topology and optical reconfigurability requirements.
  • Perform system-level simulation of end-to-end fabric performance.
  • Drive failure analysis, root cause analysis, and corrective actions for production interconnect issues.
  • Contribute to tooling and automation for health monitoring, telemetry, diagnostics, and remediation.
  • Stay current with industry standards and emerging technologies.

Basic Qualifications

  • 8+ years of hands-on experience with high-speed copper and optical interconnects.
  • Master’s or PhD degree in Electrical Engineering, Photonics, or Physics.
  • Deep knowledge of PAM4 SerDes performance, equalization, jitter, and crosstalk.
  • Operational understanding of FEC, Retimers, TIAs, and Drivers.
  • Experience with optical link budget analysis and diagnostics.
  • Expertise in transceiver components, SiPh PICs, DSPs, and failure characterization.
  • Knowledge of thermal, mechanical, power, and signal integrity constraints.
  • Knowledge of SiPh design process, reliability testing, and yield improvement.
  • Familiarity with CPO technologies and supply chain ecosystems.
  • Strong problem-solving skills in fast-paced environments.

Compensation and Benefits

$180,000 - $440,000 USD

Compensation includes base salary, equity, medical, vision, and dental coverage, 401(k), disability insurance, life insurance, and additional perks.

Similar jobs you might like

Technology

SpaceXAI

Member of Technical Staff - RL Training Framework

Senior

On-site

Palo Alto, CA

180,000 - 440,004 USD/yr

🏢 Summary: Engineering role focused on developing and optimizing reinforcement learning training infrastructure for large-scale workloads, from experimentation to production systems. The position involves improving scalability, observability, and end-to-end training performance in distributed environments. 🗂️ Requirements: Experience with large-scale distributed systems, Ability to debug and optimize system efficiency, Problem-solving across all levels of the stack, Proficiency in Python, Proficiency in Jax, Rust, or C++, Strong communication skills 📃 Skills: Python, Jax, Rust, C++, RL, LLM 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: The RL infrastructure team is looking for an engineer to help develop our RL training framework. RESPONSIBILITIES: - Design and implement the systems backing all RL workloads, from small scale ablations to production training runs - Profile, debug, and optimize end-to-end training performance - Improve scalability and observability of the RL stack BASIC QUALIFICATIONS: - Experience building, debugging, and optimizing efficiency of large-scale distributed systems - Comfortable diving into unfamiliar areas and solving problems at all levels of the stack - Proficiency in Python, Jax, Rust, and/or C++ PREFERRED SKILLS AND EXPERIENCE: - Experience with large scale LLM training infrastructure - Strong knowledge of reinforcement learning techniques - Experience with RL numerics COMPENSATION AND BENEFITS: - $180,000 - $440,000 USD - Equity package - Medical, vision, and dental coverage - 401(k) retirement plan - Short and long-term disability insurance - Life insurance - Additional discounts and perks SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Technology

SpaceXAI

Member of Technical Staff - Mid-training

Mid

On-site

Palo Alto, CA

180,000 - 440,004 USD/yr

🏢 Summary: AI engineering role focused on scaling synthetic training data, optimizing large-model training pipelines, and developing evaluation systems for advanced ML models. The position involves large-scale data processing, model distillation, and long-context data engineering in a highly technical environment. 🗂️ Requirements: Expertise in machine learning and large model scaling, Knowledge of scaling laws, Ability to design ML experiments, Experience with AI training data curation, Familiarity with text, image, audio, and video modalities, Strong engineering skills in large-scale data processing frameworks, Experience with Spark, Experience with Ray, Strong communication skills 📃 Skills: ML, Spark, Ray, Docker, RL, AI 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. RESPONSIBILITIES: - Scale synthetic coding data to trillions of tokens with large-scale docker verification. - Distill the intelligence of flagship models into flash models through synthetic data generation. - Optimize mid-training data mixtures to boost the ceiling for RL. - Engineer long-context data recipes. - Develop robust and diverse evaluation for mid-training checkpoints. BASIC QUALIFICATIONS: - Expertise in ML and large model scaling, with familiarity across all kinds of scaling laws. - Strong ability to design ML experiments. - Familiarity with state-of-the-art techniques for curating AI training data for text, image, audio, and video modalities. - Strong engineering abilities in Spark, Ray, and other frameworks for large-scale data processing. COMPENSATION AND BENEFITS: - $180,000 - $440,000 USD - Base salary is part of a total rewards package that includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and additional discounts and perks. SpaceXAI is an equal opportunity employer. For details on data processing, view the Recruitment Privacy Notice.

Technology

SpaceXAI

Software Engineer - Networking Software and Services

Senior

On-site

Palo Alto, CA

🏢 Summary: AI network software engineering role focused on building automation-first tools and services for large-scale GPU supercomputing network fabrics. The position involves developing scalable network management systems, implementing infrastructure-as-code practices, and improving deployment reliability for AI training and inference environments. 🗂️ Requirements: Deep experience collaborating with network engineers, Strong knowledge of network topologies, Strong knowledge of network protocols, Experience designing scalable and reliable software systems, Ability to orchestrate large-scale network devices, Experience with infrastructure-as-code practices, Strong communication skills, Ability to work hands-on in ambiguous environments 📃 Skills: Python, Go, TCP/IP, BGP, RDMA, IaC, Networking, Automation 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. About the Role As part of the Network Software and Services for AI (nssAI) team at xAI, you'll build cutting-edge software, services, and frameworks to empower our Network Development Engineers. Working hands-on, you'll tackle all facets of network management—metric collection, configuration, zero-touch provisioning, monitoring, and auto-remediation—driving automation-first solutions for production and ancillary networks. Expect to develop extensible tools, streamline complex processes, and ensure rock-solid reliability to support AI-driven scientific discovery. Focus - Building software and tools with extensive metrics coverage for large GPU supercomputing network fabrics used for AI training and customer inference queries. - Implementing IaC best practices, enhancing deployment pipelines, and ensuring robust, secure service delivery across production environments. Preferred Skills and Experience - Deep experience collaborating with network engineers using extensive knowledge of network topologies and network protocols. - Expert knowledge and proven history designing scalable and reliable software systems from the ground up. - Ability to build and orchestrate tens of thousands of network devices efficiently. - Ability to thrive in ambiguity and create metrics that help prioritize team focus. Tech Stack - Python - Go - TCP/IP - BGP - RDMA Annual Salary Range - $150,000 - $250,000 base Benefits - Equity package - Medical, vision, and dental coverage - 401(k) retirement plan - Short and long-term disability insurance - Life insurance - Additional discounts and perks SpaceXAI is an equal opportunity employer. For details on data processing, view the Recruitment Privacy Notice.

Technology

SpaceXAI

Member Of Technical Staff - Cloud Infrastructure

Senior

On-site

Palo Alto, CA

180,000 - 440,004 USD/yr

🏢 Summary: Senior Infrastructure Engineer role focused on building and operating secure, scalable AI infrastructure for US government projects across bare metal and classified cloud environments. The position involves Kubernetes-based infrastructure management, GPU cluster operations, observability, automation, and compliance-driven reliability engineering. This is an in-person role in Palo Alto or Washington, DC with up to 50% travel. 🗂️ Requirements: Active Top Secret (TS) security clearance, 5+ years of infrastructure or site reliability engineering experience, Experience building and maintaining scalable systems, Proficiency with Pulumi, Terraform, or Ansible, Deep knowledge of Kubernetes, CNI, CRI, and CSI, Experience with incident management and SLAs/SLOs, Strong communication and documentation skills, Ability to work in secure or government environments, Willingness to travel up to 50%, On-site availability in Palo Alto, CA or Washington, DC 📃 Skills: Kubernetes, CNI, CRI, CSI, Pulumi, Terraform, Ansible, GPU, Kyverno, ArgoCD, Go, IaC, Observability, SLA, SLO 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: We are seeking a highly skilled Senior Infrastructure Engineer to join our US Government Team, focused on designing, building, and operating secure, scalable infrastructure for critical government projects. In this role, you will develop and manage training and inference clusters, as well as highly reliable applications, across bare metal, classified cloud, and hybrid cloud architectures. You will leverage your expertise in Kubernetes and GPU hardware to deliver robust, secure systems that support large-scale AI workloads while meeting stringent federal compliance requirements. This role demands a passion for automation, observability, and ensuring system integrity in a fast-paced, high-security environment. RESPONSIBILITIES: - Develop and optimize software to provision and manage xAI's infrastructure across on-premise, virtual machine, and classified cloud environments, enabling efficient scaling for US government initiatives. - Enhance the reliability, performance, and cost-effectiveness of infrastructure to support large-scale AI and application workloads in secure, classified settings. - Collaborate with xAI engineers to understand workload requirements and design tailored solutions that meet government-specific needs and compliance standards. - Implement robust observability, monitoring, and security practices to ensure the integrity, availability, and confidentiality of critical systems, adhering to federal protocols. - Manage storage infrastructure using Infrastructure-as-Code (IaC) tools such as Pulumi, Terraform, or Ansible, with a focus on secure data handling. - Drive system reliability through incident management, postmortems, and the definition of clear SLAs and SLOs, while maintaining security and compliance. - This is an in-person role based in Palo Alto, CA or Washington, DC, with up to 50% travel required. BASIC QUALIFICATIONS: - Active Top Secret (TS) security clearance. - 5+ years of experience as an Infrastructure Engineer, Site Reliability Engineer, or similar role, with a focus on building and maintaining reliable, scalable systems, preferably in secure or government environments. - Proficiency in managing storage infrastructure with IaC tools such as Pulumi, Terraform, or Ansible. - Deep understanding of the Kubernetes stack, including CNI, CRI, CSI, and related components. - Demonstrated ability to improve system reliability through incident management, postmortems, and defining SLAs/SLOs. - Excellent communication and documentation skills, with the ability to handle sensitive information concisely and accurately. PREFERRED SKILLS AND EXPERIENCE: - Deep familiarity with installing and using GPU hardware, including setting up drivers, debugging issues, and ensuring reliability. - Experience with high-traffic web or mobile application workloads, including optimizing Kubernetes for large-scale deployments in classified or federal settings. - Familiarity with chaos engineering, capacity planning, or similar practices for ensuring system resilience in government projects. - Proficiency with tools such as Kyverno, ArgoCD, or Go programming for infrastructure automation. - Strong sense of ownership, curiosity, and enthusiasm for tackling complex technical challenges in secure environments. - Passion for problem-solving and a proactive drive to deliver impactful results while adhering to security protocols. - Certifications in security-related fields (e.g., CISSP) or experience in secure federal environments. COMPENSATION AND BENEFITS: - $180,000 - $440,000 USD - Base salary is just one part of the total rewards package, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks. SpaceXAI is an equal opportunity employer. For details on data processing, view the Recruitment Privacy Notice.

Technology

xAI

Software Engineer - Network (C++)

Mid

On-site

Seattle, WA

🏢 Summary: Software Engineer role focused on developing and operating high-performance networking software for large-scale AI datacenter infrastructure powering GPU clusters and frontier AI models. The position involves building routing algorithms, real-time distributed systems, deployment tooling, and reliable low-latency networking software at massive scale. Engineers own the full software lifecycle from architecture and implementation to deployment and monitoring. 🗂️ Requirements: Bachelor's degree in computer science, engineering, math, or related technical discipline OR 2+ years of professional software development experience, Strong development experience in C or C++, Experience developing and deploying software at scale, Knowledge of networking protocols, Knowledge of distributed systems, Ability to work in fast-paced environments, Strong written and verbal communication skills, Willingness to work extended hours and weekends as needed 📃 Skills: C, C++, UDP, TCP/IP, RDMA, Networking, DistributedSystems, HPC, Linux, CI, Monitoring, Visualization, Virtualization 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: At SpaceXAI, we design, build, and operate Colossus from the ground up. This includes the massive GPU clusters, high-speed interconnect fabric, and the software that makes it all work at unprecedented scale. Colossus powers Grok and our frontier AI models with a custom, high-performance datacenter network that delivers ultra-low latency and massive bandwidth across hundreds of thousands of GPUs. As a Software Engineer on the Colossus Networking team, you will develop the core networking software that maximizes the performance and reliability of our datacenter fabric. Your work will directly impact training efficiency, model convergence, and the speed at which we can push the frontier of AI. Our engineers own the full lifecycle of their software — from design and implementation to deployment, monitoring, and iteration based on real-world performance at scale. You will solve hard problems in distributed systems, high-performance networking, and real-time control of one of the largest AI supercomputers on Earth. RESPONSIBILITIES: - Develop routing and traffic-engineering algorithms for the Colossus high-performance datacenter network. - Develop highly reliable, real-time software designed to run on the switches that form the backbone of our low-latency, high-bandwidth AI training fabric. - Participate in and lead architecture, design, and code reviews. - Develop prototypes and run experiments to validate key design decisions at both small and full-cluster scale. - Build tools for software development, deployment, data analysis, visualization, and testing across virtualized environments, hardware-in-the-loop setups, and live production clusters. - Deploy reliable software updates through continuous integration and release systems with rigorous testing and monitoring. BASIC QUALIFICATIONS: - Bachelor's degree in computer science, engineering, math, or a related technical discipline; OR 2+ years of professional software development experience in lieu of a degree. - Strong development experience in C or C++. PREFERRED SKILLS AND EXPERIENCE: - Strong professional experience writing high-performance C/C++ in production environments. - Experience developing, debugging, and deploying software that runs at scale in real-world systems. - Deep knowledge of networking protocols (UDP, TCP/IP, RDMA, etc.), distributed systems, and large-scale datacenter fabrics. - Background in real-time systems, high-performance computing, low-latency networking, or resource-constrained environments. - Creative problem-solving ability with exceptional analytical skills and strong engineering fundamentals. - Excellent written and verbal communication skills. - Ability to thrive in a fast-paced, dynamic environment with evolving requirements. - Experience with security considerations in large-scale distributed systems. ADDITIONAL REQUIREMENTS: - Must be willing to work extended hours and weekends as needed. COMPENSATION AND BENEFITS - $180,000 - $440,000 USD - Equity package - Comprehensive medical, vision, and dental coverage - 401(k) retirement plan - Short and long-term disability insurance - Life insurance - Additional discounts and perks SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Technology

xAI

Software Engineer - Network (C++)

Mid

On-site

Palo Alto, CA

🏢 Summary: Software Engineer role focused on building and operating high-performance networking software for large-scale AI datacenter infrastructure powering GPU clusters and frontier AI models. The position involves developing routing, traffic-engineering, and real-time distributed systems software with ownership across the full software lifecycle. Engineers work on low-latency networking, deployment, monitoring, and scalability challenges in production environments. 🗂️ Requirements: Bachelor's degree in computer science, engineering, math, or related discipline OR 2+ years of professional software development experience, Strong development experience in C or C++, Experience developing and deploying software at scale, Knowledge of networking protocols and distributed systems, Ability to work in fast-paced environments with evolving requirements, Strong written and verbal communication skills, Willingness to work extended hours and weekends when needed 📃 Skills: C, C++, UDP, TCP/IP, RDMA, Networking, DistributedSystems, Linux, CI, HPC 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: At SpaceXAI, we design, build, and operate Colossus from the ground up. This includes the massive GPU clusters, high-speed interconnect fabric, and the software that makes it all work at unprecedented scale. Colossus powers Grok and our frontier AI models with a custom, high-performance datacenter network that delivers ultra-low latency and massive bandwidth across hundreds of thousands of GPUs. As a Software Engineer on the Colossus Networking team, you will develop the core networking software that maximizes the performance and reliability of our datacenter fabric. Your work will directly impact training efficiency, model convergence, and the speed at which we can push the frontier of AI. Our engineers own the full lifecycle of their software — from design and implementation to deployment, monitoring, and iteration based on real-world performance at scale. You will solve hard problems in distributed systems, high-performance networking, and real-time control of one of the largest AI supercomputers on Earth. RESPONSIBILITIES: - Develop routing and traffic-engineering algorithms for the Colossus high-performance datacenter network. - Develop highly reliable, real-time software designed to run on the switches that form the backbone of our low-latency, high-bandwidth AI training fabric. - Participate in and lead architecture, design, and code reviews. - Develop prototypes and run experiments to validate key design decisions at both small and full-cluster scale. - Build tools for software development, deployment, data analysis, visualization, and testing across virtualized environments, hardware-in-the-loop setups, and live production clusters. - Deploy reliable software updates through continuous integration and release systems with rigorous testing and monitoring. BASIC QUALIFICATIONS: - Bachelor's degree in computer science, engineering, math, or a related technical discipline; OR 2+ years of professional software development experience in lieu of a degree. - Strong development experience in C or C++. PREFERRED SKILLS AND EXPERIENCE: - Strong professional experience writing high-performance C/C++ in production environments. - Experience developing, debugging, and deploying software that runs at scale in real-world systems. - Deep knowledge of networking protocols (UDP, TCP/IP, RDMA, etc.), distributed systems, and large-scale datacenter fabrics. - Background in real-time systems, high-performance computing, low-latency networking, or resource-constrained environments. - Creative problem-solving ability with exceptional analytical skills and strong engineering fundamentals. - Excellent written and verbal communication skills. - Ability to thrive in a fast-paced, dynamic environment with evolving requirements. - Experience with security considerations in large-scale distributed systems. ADDITIONAL REQUIREMENTS: - Must be willing to work extended hours and weekends as needed. COMPENSATION AND BENEFITS - $180,000 - $440,000 USD - Base salary is just one part of the total rewards package, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks. SpaceXAI is an equal opportunity employer. For details on data processing, view the Recruitment Privacy Notice.

Technology

SpaceXAI

Member of Technical Staff - Media

Senior

On-site

Seattle, WA

180,000 - 440,004 USD/yr

🏢 Summary: Media engineering role focused on building and optimizing large-scale video services integrated with advanced AI infrastructure for a platform serving hundreds of millions of users. The position involves rebuilding media processing and distribution pipelines, improving video quality and performance, and working with scalable distributed systems using high-performance languages. 🗂️ Requirements: 5+ years of experience, Proficiency in C++ or Go, Experience with WebRTC, LL-HLS, or video transcoding pipelines, Experience building scalable distributed systems, Strong focus on media quality and performance, Strong communication skills 📃 Skills: Go, C++, Rust, Java, Scala, Kubernetes, FoundationDB, ValKey, Envoy, S3, WebRTC, LL-HLS, H.264, H.265, AV1, MP4, CMAF, VMAF, RTP, RTMP, HDR, DRM 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: We're looking for exceptional media engineers who want to join us on a new project to deeply integrate xAI's advanced AI infrastructure into a platform used by around 600 million users every month. We're bringing xAI's technology stack and using it to transform the video product experience - video playback, live streaming, Spaces, audio/video calls, and more. This is your chance to contribute in a major way while leveraging all of the powerful AI tools and talented colleagues at xAI. RESPONSIBILITIES: - Build the next generation of large-scale video services - Contribute to and rebuild core media processing and distribution pipelines in high-performance languages (Rust, C++ or Go) - Obsess over every millisecond and pixel, ensuring end-to-end media quality and performance at scale across a rich suite of products and user platforms BASIC QUALIFICATIONS: - At least 5 years of experience - Obsessed with media quality, performance, and product experience - Proficient in high performance C++ or Go - In-depth knowledge of either WebRTC or LL-HLS or video transcoding pipelines - Familiar with building and running scalable and resilient distributed systems PREFERRED SKILLS AND EXPERIENCE: - Go, C++, Rust, Java, Scala - Kubernetes, FoundationDB, ValKey, Envoy, S3 - H.264, H.265, AV1, MP4, CMAF, VMAF, RTP, RTMP, LL-HLS, HDR, DRM COMPENSATION AND BENEFITS: $180,000 - $440,000 USD Base salary is just one part of our total rewards package, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks. SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Technology

xAI

Software Engineer - Training/Inference (C++)

Senior

On-site

Palo Alto, CA

🏢 Summary: High-impact inference engineering role focused on building and optimizing large-scale distributed model serving systems for Grok. The position involves low-level GPU and inference optimization, scalable infrastructure development, and ensuring high-performance, reliable AI serving at massive scale. Candidates will work across the full inference stack, from orchestration to GPU kernels and CI/CD systems. 🗂️ Requirements: Deep systems programming in C/C++ or Rust, Experience with large-scale high-concurrency production serving, Experience with GPU inference engines, Strong knowledge of batching, caching, load balancing, and parallelism, Experience with GPU kernel optimization and code generation, Knowledge of quantization, speculative decoding, distillation, and low-precision numerics, Experience with testing, benchmarking, and reliability engineering for inference services, Experience designing and implementing CI/CD infrastructure for inference 📃 Skills: C, C++, Rust, vLLM, SGLang, Triton, TensorRT-LLM, GPU, CI/CD, Quantization, Distillation, Parallelism, Caching, Benchmarking, Load-balancing, Autoscaling 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. About the Role: We are building the high-performance inference platform that serves Grok to millions of users every day with lightning speed and perfect reliability. As a Member of Technical Staff - Inference, you will design and optimize large-scale model serving systems end-to-end. You will own everything from distributed infrastructure (global KV cache, continuous batching, load balancing, auto-scaling) to deep low-level optimizations (GPU kernels, quantization, speculative decoding, tail latency). This is a high-impact role where your work directly determines how fast and reliably users interact with Grok at massive scale. Responsibilities: - Architect and implement scalable distributed infrastructure for model serving (load balancing, auto-scaling, batch scheduling, global KV cache). - Optimize latency and throughput of model inference under real production workloads. - Build reliable, high-concurrency serving systems that serve billions of users with 100% uptime, 0% error rate, and excellent tail latency. - Benchmark, fine-tune, and accelerate inference engines (including low-level GPU kernel work and code generation). - Develop custom tools to trace, replay, and fix issues across the full stack — from orchestration down to GPU kernels. - Create robust CI/CD infrastructure for seamless endpoint deployment, image publishing, and inference engine updates. - Accelerate research on scaling test-time compute, RL rollout, and model-hardware co-design for next-generation systems. Basic Qualifications: - Deep low-level systems programming (C/C++ or Rust) - Experience with large-scale, high-concurrent production serving. - Experience with GPU inference engines (vLLM, SGLang, Triton, TensorRT-LLM, etc.). - Strong background in system optimizations: batching, caching, load balancing, parallelism. - Low-level inference optimizations: GPU kernels, code generation. - Algorithmic inference optimizations: quantization, speculative decoding, distillation, low-precision numerics. - Experience with testing, benchmarking, and reliability of inference services. - Experience designing and implementing CI/CD infrastructure for inference. Compensation and Benefits: - $180,000 - $440,000 USD - Base salary is just one part of the total rewards package, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks. SpaceXAI is an equal opportunity employer. For details on data processing, view the Recruitment Privacy Notice.

Technology

SpaceXAI

Member of Technical Staff - Media

Senior

On-site

Palo Alto, CA

180,000 - 440,004 USD/yr

🏢 Summary: Media engineering role focused on building and optimizing large-scale video services integrated with advanced AI infrastructure for a platform serving hundreds of millions of users. The position involves rebuilding media processing and distribution pipelines using high-performance languages and improving video quality, streaming, and real-time communication systems at scale. Candidates should have strong distributed systems experience and deep expertise in modern media technologies. 🗂️ Requirements: 5+ years of experience, Proficiency in C++ or Go, Knowledge of WebRTC, LL-HLS, or video transcoding pipelines, Experience with scalable distributed systems, Strong media quality and performance optimization skills, Strong communication skills 📃 Skills: Go, C++, Rust, Java, Scala, Kubernetes, FoundationDB, ValKey, Envoy, S3, WebRTC, LL-HLS, H.264, H.265, AV1, MP4, CMAF, VMAF, RTP, RTMP, HDR, DRM 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: We're looking for exceptional media engineers who want to join us on a new project to deeply integrate xAI's advanced AI infrastructure into a platform used by around 600 million users every month. We're bringing xAI's technology stack and using it to transform the video product experience - video playback, live streaming, Spaces, audio/video calls, and more. This is your chance to contribute in a major way while leveraging all of the powerful AI tools and talented colleagues at xAI. RESPONSIBILITIES: - Build the next generation of large-scale video services - Contribute to and rebuild core media processing and distribution pipelines in high-performance languages (Rust, C++ or Go) - Obsess over every millisecond and pixel, ensuring end-to-end media quality and performance at scale across a rich suite of products and user platforms BASIC QUALIFICATIONS: - At least 5 years of experience - Obsessed with media quality, performance, and product experience - Proficient in high performance C++ or Go - In-depth knowledge of either WebRTC or LL-HLS or video transcoding pipelines - Familiar with building and running scalable and resilient distributed systems PREFERRED SKILLS AND EXPERIENCE: - Go, C++, Rust, Java, Scala - Kubernetes, FoundationDB, ValKey, Envoy, S3 - H.264, H.265, AV1, MP4, CMAF, VMAF, RTP, RTMP, LL-HLS, HDR, DRM COMPENSATION AND BENEFITS: - $180,000 - $440,000 USD - Base salary is just one part of the total rewards package, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks. Equal opportunity employer. Recruitment Privacy Notice available.

Technology

xAI

Member of Technical Staff - RL Training Framework

Senior

On-site

Palo Alto, CA

180,000 - 440,004 USD/yr

🏢 Summary: Engineering role focused on developing and optimizing reinforcement learning training infrastructure for large-scale and production AI workloads. The position involves building distributed systems, improving scalability and observability, and enhancing end-to-end training performance. Candidates should have strong systems engineering skills and experience with modern AI training infrastructure. 🗂️ Requirements: Experience building large-scale distributed systems, Experience debugging and optimizing system efficiency, Ability to solve problems across all levels of the stack, Proficiency in Python, Proficiency in Jax, Proficiency in Rust or C++, Strong communication skills 📃 Skills: Python, Jax, Rust, C++, RL, LLM 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: The RL infrastructure team is looking for an engineer to help develop our RL training framework. RESPONSIBILITIES: - Design and implement the systems backing all RL workloads at SpaceXAI, from small scale ablations to production training runs. - Profile, debug, and optimize end-to-end training performance - Improve scalability and observability of the RL stack BASIC QUALIFICATIONS: - Experience building, debugging, and optimizing efficiency of large-scale distributed systems - Comfortable diving into unfamiliar areas and solving problems at all levels of the stack - Proficiency in Python, Jax, Rust, and/or C++ PREFERRED SKILLS AND EXPERIENCE: - Experience with large scale LLM training infrastructure - Strong knowledge of reinforcement learning techniques - Experience with RL numerics COMPENSATION AND BENEFITS: $180,000 - $440,000 USD Base salary is just one part of our total rewards package, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks. SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.