June 30, 2026

Backend Engineer - API

Senior • On-site

Palo Alto, CA

About the Role

As an ideal candidate you have a good understanding of how highly scalable and reliable production infrastructure is built. Most of the backend infrastructure is written in Rust. Familiarity with a compiled language such as C++, Rust, or Go is highly beneficial.

Responsibilities

  • Build the xAI API that serves models to developers worldwide
  • Own the end-to-end system responsible for high-throughput inference, handling billions of tokens per minute with low latency and high availability, including model serving infrastructure, request routing, SDK development, rate limiting, observability, and efficient scaling

Basic Qualifications

  • Expert knowledge of either Rust or C++
  • Experience in designing, implementing, and maintaining reliable and horizontally scalable distributed systems
  • Knowledge of service observability and reliability best practices
  • Experience in operating commonly used databases such as PostgreSQL, Clickhouse, and MongoDB

Preferred Skills and Experience

  • Experience with LLM inference engines and serving frameworks (e.g., SGLang, TensorRT, vLLM)
  • Experience designing or building with agent SDKs and agent orchestration frameworks
  • Experience with Docker, Kubernetes, and containerized applications
  • Expert knowledge of gRPC (unary, response streaming, bi-directional streaming, REST mapping)

Compensation and Benefits

$180,000 - $440,000 USD

Base salary is just one part of the total rewards package, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short- and long-term disability insurance, life insurance, and various other discounts and perks.

xAI is an equal opportunity employer. For details on data processing, view the Recruitment Privacy Notice.

Similar jobs you might like

Technology

xAI

Backend Engineer - API

Senior

On-site

Palo Alto, CA

🏢 Summary: Engineering role focused on building and owning a high-throughput, low-latency API and backend infrastructure for large-scale model inference. The position involves designing and operating reliable, horizontally scalable distributed systems that serve billions of tokens per minute. You will develop and maintain model serving, routing, SDKs, and observability within a production-grade environment. 🗂️ Requirements: Expert knowledge of Rust or C++, Experience building and maintaining horizontally scalable distributed systems, Experience designing reliable high-availability production infrastructure, Knowledge of observability and reliability best practices, Experience operating PostgreSQL, Clickhouse, or MongoDB 📃 Skills: Rust, C++, Go, PostgreSQL, Clickhouse, MongoDB, gRPC, Docker, Kubernetes, TensorRT, vLLM, SGLang, REST, SDK 🏢 Description: About xAI xAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.ABOUT THE ROLE: As an ideal candidate you have a good understanding of how highly scalable and reliable production infrastructure is built. Most of our backend infrastructure is written in Rust. So familiarity with a compiled language such as C++, Rust, or Go is highly beneficial. RESPONSIBILITIES: Build the xAI API that serves our models to developers worldwide Own the end-to-end system responsible for high-throughput inference, handling billions of tokens per minute with low latency and high availability, including model serving infrastructure, request routing, SDK development, rate limiting, observability, and efficient scaling BASIC QUALIFICATIONS: Expert knowledge of either Rust or C++ Experience in designing, implementing, and maintaining reliable and horizontally scalable distributed systems Knowledge of service observability and reliability best practices Experience in operating commonly used databases such as PostgreSQL, Clickhouse, and MongoDB PREFERRED SKILLS AND EXPERIENCE: Experience with LLM inference engines and serving frameworks (e.g., SGLang, TensorRT, vLLM) Experience designing or building with agent SDKs and agent orchestration frameworks Experience with Docker, Kubernetes, and containerized applications Expert knowledge of gRPC (unary, response streaming, bi-directional streaming, REST mapping) COMPENSATION AND BENEFITS $180,000 - $440,000 USD Base salary is just one part of our total rewards package at xAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.xAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Technology

xAI

Exceptional Software Engineer

Senior

On-site

Palo Alto, CA

🏢 Summary: This role involves working on the most critical product or technical challenges, contributing directly to the development of truth-seeking AI systems. The position is hands-on and impact-driven, with clear project ownership defined before the offer stage. It offers a highly competitive compensation package including equity and comprehensive benefits. 🗂️ Requirements: Strong belief in truth-seeking AI as a critical problem, Exceptional software engineering skills, Ability to thrive in meritocratic environments, High work ethic and strong prioritization skills, Strong communication skills, Ability to contribute hands-on to critical technical challenges 📃 Skills: Software, Engineering, AI 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.ABOUT THE ROLE: You will work on the most critical product or technical challenge at any given time. You will get clarity on your first project before an offer. BASIC QUALIFICATIONS: You believe truth-seeking AI is the most important and challenging problem. You take pride in your work and thrive in meritocratic environments. You are an exceptional software engineer. COMPENSATION AND BENEFITS: $180,000 - $440,000 USD Base salary is just one part of our total rewards package at SpaceXAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks. SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Technology

xAI

Software Engineer - Kernels/CUDA (C++)

Senior

On-site

Seattle, WA

🏢 Summary: High-impact compute infrastructure role focused on building and optimizing massive GPU supercomputers for AI training and inference. The position involves low-level systems programming, GPU kernel optimization, Linux kernel internals, orchestration, and distributed infrastructure to improve scalability, reliability, and performance of AI workloads. 🗂️ Requirements: Deep systems programming experience in C/C++ or Rust, Experience with large-scale GPU clusters or distributed compute infrastructure, Hands-on GPU kernel optimization using CUTLASS, custom kernels, or Nsight, Knowledge of Linux kernel internals, scheduling, virtualization, or orchestration, Experience building high-performance AI training or inference infrastructure, Ability to optimize memory-bound and compute-bound workloads, Experience operating exabyte-scale storage systems 📃 Skills: CUDA, CUTLASS, Tensor, Nsight, Linux, KVM, Firecracker, Kubernetes, C++, Rust 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.ABOUT THE ROLE: We are building one of the world's largest AI supercomputers from the ground up. As part of the Compute Infrastructure team, you will own both the raw GPU supercomputer and the platform layer that runs on top of it. You will work across the full stack — from low-level GPU kernel optimizations and Linux kernel internals to massive-scale orchestration and virtualization — to make training and inference at SpaceXAI as fast, reliable, and scalable as possible. This is a broad, high-impact role that combines hardcore supercompute and compute infrastructure work. Your contributions will directly accelerate Grok's training speed and overall AI progress. RESPONSIBILITIES: Design, build, and optimize massive GPU clusters for extreme-scale training and inference workloads Develop and tune low-level CUDA kernels (GeMM, Attention, etc.), using CUTLASS, Tensor Cores, and Nsight for maximum performance Profile, debug, and eliminate bottlenecks across GPU memory hierarchy, networking fabric, filesystems, and multi-GPU operation Collaborate closely with AI research teams to deliver production-grade performance and scalability PREFERRED SKILLS AND EXPERIENCE: Deep low-level systems programming (C/C++/PTX/SASS) Strong experience with large-scale GPU clusters or distributed compute infrastructure at production scale Hands-on work with GPU kernel optimization (CUTLASS, custom kernels, Nsight profiling) Track record of building or running high-performance infrastructure for AI workloads (training or inference platforms) Ability to reason from first principles and optimize for both memory-bound and compute-bound scenarios COMPENSATION AND BENEFITS: $180,000 - $440,000 USD Base salary is just one part of our total rewards package at SpaceXAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Technology

xAI

Exceptional Software Engineer

Senior

On-site

Seattle, WA

🏢 Summary: This role involves working on the most critical product or technical challenges within a cutting-edge AI organization focused on building truth-seeking AI systems. The position requires direct, hands-on contribution to high-impact engineering problems in a fast-paced, meritocratic environment. Candidates will tackle core technical initiatives aligned with advancing AI capabilities. 🗂️ Requirements: Exceptional software engineering skills, Proven ability to solve complex technical problems, Hands-on experience in building and delivering software systems, Ability to work on high-priority product or technical challenges, Strong technical communication skills 📃 Skills: Software, Engineering, AI, Programming, Systems 🏢 Description: About xAI xAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.ABOUT THE ROLE: You will work on the most critical product or technical challenge at any given time. You will get clarity on your first project before an offer. BASIC QUALIFICATIONS: You believe truth-seeking AI is the most important and challenging problem. You take pride in your work and thrive in meritocratic environments. You are an exceptional software engineer. COMPENSATION AND BENEFITS: $180,000 - $440,000 USD Base salary is just one part of our total rewards package at xAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks. xAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Technology

xAI

Exceptional Software Engineer

Senior

On-site

Palo Alto, CA

🏢 Summary: Opportunity for an exceptional software engineer to tackle the most critical product or technical challenges within a high-impact AI environment. The role involves hands-on contribution to core systems and direct influence on cutting-edge AI development. 🗂️ Requirements: Exceptional software engineering ability, Proven experience building and delivering complex software systems, Ability to work on high-priority technical challenges, Hands-on development experience in production environments, Strong technical communication skills 📃 Skills: Programming, Algorithms, DataStructures, Systems, AI, Testing, Debugging 🏢 Description: About xAI xAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.ABOUT THE ROLE: You will work on the most critical product or technical challenge at any given time. You will get clarity on your first project before an offer. BASIC QUALIFICATIONS: You believe truth-seeking AI is the most important and challenging problem. You take pride in your work and thrive in meritocratic environments. You are an exceptional software engineer. COMPENSATION AND BENEFITS: $180,000 - $440,000 USD Base salary is just one part of our total rewards package at xAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks. xAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Technology

xAI

Sr. Software Engineer (Data Center Automation)

Senior

On-site

Palo Alto, CA

🏢 Summary: Senior Software Engineer role focused on building automation and observability solutions to enhance reliability across multi-data center AI infrastructure. The position combines strong programming skills with hands-on data center and Linux systems expertise to minimize downtime and optimize performance. It involves developing scalable services, improving monitoring and incident response, and collaborating across infrastructure and facility teams. 🗂️ Requirements: Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering or related field (or equivalent experience), 3+ years experience in SRE, Infrastructure, DevOps, or Systems Engineering in large-scale production environments, Strong production experience in Python, Solid Linux systems administration and kernel-level knowledge, Experience with containerization and orchestration (Docker, Kubernetes or similar), Experience implementing observability solutions (metrics, logging, tracing, monitoring, alerting), Understanding of networking fundamentals (TCP/IP, routing, DNS, redundancy), Experience troubleshooting distributed systems, hardware and network issues, Experience with on-call rotations and incident response practices (SLAs, error budgets), Ability to collaborate with cross-functional technical teams 📃 Skills: Python, Rust, Linux, Kubernetes, Docker, Prometheus, Grafana, TCP/IP, DNS, Scripting, Automation, Observability, Monitoring, Tracing, Networking 🏢 Description: ABOUT xAI xAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: We are seeking a highly skilled Sr. Software Engineer to join our team in managing and enhancing reliability across a multi-data center environment. This role focuses on automating processes, building and implementing robust observability solutions, and ensuring seamless operations for mission-critical AI infrastructure. The ideal candidate will combine strong coding abilities with hands-on data center experience to build scalable reliability services, optimize system performance, and minimize downtime—including close partnership with facility operations to address physical infrastructure impacts. In an era where AI workloads demand near-zero downtime, this position plays a pivotal role in bridging software engineering principles with physical data center realities. By prioritizing automation and observability, team members in this role can reduce mean time to recovery (MTTR) by up to 50% through proactive monitoring and automated remediation. The primary objective of this team is to mitigate downtime and minimize impact to end-users from both scheduled and unscheduled maintenance, as well as events affecting onsite data centers. This is achieved through proactive automation, robust observability, and integrated software-physical reliability strategies. RESPONSIBILITIES: - Design, develop, and deploy scalable code and services (primarily in Python and Rust) to automate reliability workflows, including monitoring, alerting, incident response, and infrastructure provisioning. - Implement and maintain observability tools and practices, such as metrics collection, logging, tracing, and dashboards, to provide real-time insights into system health across multiple data centers. - Collaborate with cross-functional teams to identify reliability bottlenecks and automate solutions for fault tolerance, disaster recovery, capacity planning, and physical/environmental risk mitigation. - Troubleshoot and resolve complex issues in data center environments, including hardware failures, environmental anomalies, software bugs, and network-related problems, while adhering to reliability principles like error budgets and SLAs. - Optimize Linux-based systems for performance, security, and reliability, including kernel tuning, container orchestration, and scripting for automation. - Understand network topologies and concepts in large-scale, multi-data center environments to troubleshoot connectivity, routing, redundancy, and performance issues. - Participate in on-call rotations, post-incident reviews (blameless postmortems), and continuous improvement initiatives. - Mentor junior team members and document processes to foster a culture of automation and knowledge sharing. BASIC QUALIFICATIONS: - Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or a closely related technical field (or equivalent professional experience). - 3+ years of hands-on experience in site reliability engineering (SRE), infrastructure engineering, DevOps, or systems engineering in large-scale, distributed, or production environments. - Strong programming skills with proven production experience in Python; experience with Rust or another systems-level language (e.g., Go, C++) is essential. - Solid experience with Linux systems administration, performance tuning, kernel-level understanding, and scripting/automation in production environments. - Practical knowledge of containerization and orchestration technologies, such as Docker and Kubernetes. - Experience implementing observability solutions, including metrics, logging, tracing, monitoring tools, alerting, and dashboards. - Familiarity with troubleshooting complex issues in distributed systems, including software bugs, hardware failures, network problems, and environmental factors. - Understanding of networking fundamentals (TCP/IP, routing, redundancy, DNS) in large-scale or multi-site environments. - Experience participating in on-call rotations, incident response, post-incident reviews, and reliability practices such as error budgets or SLAs. - Ability to collaborate effectively with cross-functional teams. PREFERRED SKILLS AND EXPERIENCE: - 5+ years of experience in SRE or infrastructure roles in hyperscale, cloud, or AI/ML training environments with multi-data center setups. - Hands-on experience operating or scaling Kubernetes clusters at large scale, including automation for provisioning and high availability. - Proficiency in Rust for systems programming and performance-critical components. - Experience integrating software reliability tools with physical data center infrastructure (power, cooling, environmental monitoring). - Experience building automated remediation, fault tolerance, disaster recovery, capacity planning, or predictive failure detection systems. - Background in optimizing Linux-based systems for AI workloads, GPU clusters, or high-throughput compute environments. - Experience with bare-metal provisioning, data center interconnects, or hybrid/multi-site failover mechanisms. - Mentoring experience and strong documentation skills. xAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Technology

xAI

Software Engineer - Kernels/CUDA (C++)

Senior

On-site

Palo Alto, CA

🏢 Summary: High-impact compute infrastructure role focused on building and optimizing one of the world’s largest AI supercomputers for large-scale training and inference workloads. The position involves low-level GPU optimization, Linux kernel and virtualization work, distributed systems, and orchestration across massive GPU clusters. Engineers will collaborate closely with AI research teams to maximize performance, scalability, and reliability. 🗂️ Requirements: Deep systems programming experience in C/C++ or Rust, Experience with large-scale GPU clusters or distributed compute infrastructure, Hands-on GPU kernel optimization experience, Knowledge of Linux kernel internals, scheduling, virtualization, or orchestration, Experience building high-performance AI infrastructure for training or inference, Ability to optimize memory-bound and compute-bound workloads, Experience with exabyte-scale storage systems 📃 Skills: CUDA, CUTLASS, Tensor, Nsight, Linux, KVM, Firecracker, Kubernetes, C++, Rust, GPU, CUDA, GeMM, Attention 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.ABOUT THE ROLE: We are building one of the world's largest AI supercomputers from the ground up. As part of the Compute Infrastructure team, you will own both the raw GPU supercomputer and the platform layer that runs on top of it. You will work across the full stack — from low-level GPU kernel optimizations and Linux kernel internals to massive-scale orchestration and virtualization — to make training and inference at SpaceXAI as fast, reliable, and scalable as possible. This is a broad, high-impact role that combines hardcore supercompute and compute infrastructure work. Your contributions will directly accelerate Grok's training speed and overall AI progress. RESPONSIBILITIES: Design, build, and optimize massive GPU clusters for extreme-scale training and inference workloads Develop and tune low-level CUDA kernels (GeMM, Attention, etc.), using CUTLASS, Tensor Cores, and Nsight for maximum performance Profile, debug, and eliminate bottlenecks across GPU memory hierarchy, networking fabric, filesystems, and multi-GPU operation Collaborate closely with AI research teams to deliver production-grade performance and scalability PREFERRED SKILLS AND EXPERIENCE: Deep low-level systems programming (C/C++/PTX/SASS) Strong experience with large-scale GPU clusters or distributed compute infrastructure at production scale Hands-on work with GPU kernel optimization (CUTLASS, custom kernels, Nsight profiling) Track record of building or running high-performance infrastructure for AI workloads (training or inference platforms) Ability to reason from first principles and optimize for both memory-bound and compute-bound scenarios COMPENSATION AND BENEFITS: $180,000 - $440,000 USD Base salary is just one part of our total rewards package at SpaceXAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Technology

xAI

Site Reliability Engineer - Cybersecurity

Senior

On-site

Palo Alto, CA

🏢 Summary: Cybersecurity / SRE role focused on securing and maintaining the reliability of a large-scale fintech platform operating in hybrid cloud environments. The position emphasizes Kubernetes and container security, SIEM management, CI/CD protection, and automation using Python and infrastructure-as-code tools. Candidates will work on mission-critical distributed systems, ensuring regulatory compliance and resilient security operations at scale. 🗂️ Requirements: Experience securing hybrid AWS/on-premises environments, Strong proficiency in Python, Strong proficiency in Terraform, Strong proficiency in Puppet, Deep expertise in Kubernetes, Experience with container security, Hands-on experience with GitHub Actions, Experience with Prometheus, Experience with Grafana, Experience with CloudWatch, Experience with Karma, Experience managing and integrating Wazuh, Experience with security scanning tools (Semgrep, Trivy, Falco), Experience with IAM and security posture management, Ability to comply with PCI and NIST CSF standards, Located in SF Bay Area or willing to relocate 📃 Skills: AWS, IAM, Python, Terraform, Puppet, Kubernetes, Docker, GitHub, Prometheus, Grafana, CloudWatch, Karma, Wazuh, Semgrep, Trivy, Falco, PCI, NIST, CI/CD 🏢 Description: ABOUT xAI xAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: The Cybersecurity / SRE team is focused on ensuring the security and reliability of X Money. This role will primarily focus on the X Money platform but will also cross over with the X Social platform. The ideal candidate will have experience in the banking, money transmission, and P2P payments industry. We emphasize working with large distributed systems and security platforms at scale, with an automation-first mindset. You'll be responsible for securing and maintaining the reliability of X Money's infrastructure. You'll work closely with cross-functional teams to enhance security measures, improve system resilience, and implement best practices. RESPONSIBILITIES: - Build and secure mission-critical applications in a hybrid cloud environment. - Manage identities and roles effectively. - Monitor and remediate infrastructure to comply with regulations and best practices (e.g., PCI, NIST CSF). - Maintain a SIEM and all data pipelines needed for reliable alerting. - Design and implement secure container standards and automation to enable frictionless developer workflows. - Maintain Kubernetes security aligned with current best practices. - Build, deploy, and maintain security operations infrastructure using Python, Terraform, and Puppet. - Secure and enhance CI/CD pipelines. - Integrate and maintain code scanning platforms. - Develop dashboards and alerts from security metrics. - Own security projects: identify issues and implement solutions. - Apply critical analysis and problem-solving skills. BASIC QUALIFICATIONS: - Proven experience securing hybrid AWS/on-premises environments, including IAM and overall security posture. - Strong proficiency in Python, Terraform, and Puppet. - Certifications like CISA, CRISC, CGEIT, Security+, CASP+, or similar preferred. - Deep expertise in Kubernetes and container security. - Hands-on expertise building GitHub Actions and workflows. - Extensive experience with Prometheus, Grafana, CloudWatch, and Karma. - Well versed in management and integrations of Wazuh. - Hands-on experience with security scanning tools (Semgrep, Trivy, Falco). - Proactive mindset with strong ownership and problem-solving skills. - Excellent critical thinking and analytical abilities. - Located in the SF Bay Area or willing to relocate. COMPENSATION AND BENEFITS: $180,000 - $440,000 USD Base salary is just one part of our total rewards package, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks. xAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Technology

SpaceXAI

Member of Technical Staff - Mid-training

Mid

On-site

Palo Alto, CA

180,000 - 440,004 USD/yr

🏢 Summary: AI engineering role focused on scaling synthetic training data, optimizing large-model training pipelines, and developing evaluation systems for advanced ML models. The position involves large-scale data processing, model distillation, and long-context data engineering in a highly technical environment. 🗂️ Requirements: Expertise in machine learning and large model scaling, Knowledge of scaling laws, Ability to design ML experiments, Experience with AI training data curation, Familiarity with text, image, audio, and video modalities, Strong engineering skills in large-scale data processing frameworks, Experience with Spark, Experience with Ray, Strong communication skills 📃 Skills: ML, Spark, Ray, Docker, RL, AI 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. RESPONSIBILITIES: - Scale synthetic coding data to trillions of tokens with large-scale docker verification. - Distill the intelligence of flagship models into flash models through synthetic data generation. - Optimize mid-training data mixtures to boost the ceiling for RL. - Engineer long-context data recipes. - Develop robust and diverse evaluation for mid-training checkpoints. BASIC QUALIFICATIONS: - Expertise in ML and large model scaling, with familiarity across all kinds of scaling laws. - Strong ability to design ML experiments. - Familiarity with state-of-the-art techniques for curating AI training data for text, image, audio, and video modalities. - Strong engineering abilities in Spark, Ray, and other frameworks for large-scale data processing. COMPENSATION AND BENEFITS: - $180,000 - $440,000 USD - Base salary is part of a total rewards package that includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and additional discounts and perks. SpaceXAI is an equal opportunity employer. For details on data processing, view the Recruitment Privacy Notice.

Technology

SpaceXAI

Member of Technical Staff - RL Training Framework

Senior

On-site

Palo Alto, CA

180,000 - 440,004 USD/yr

🏢 Summary: Engineering role focused on developing and optimizing reinforcement learning training infrastructure for large-scale workloads, from experimentation to production systems. The position involves improving scalability, observability, and end-to-end training performance in distributed environments. 🗂️ Requirements: Experience with large-scale distributed systems, Ability to debug and optimize system efficiency, Problem-solving across all levels of the stack, Proficiency in Python, Proficiency in Jax, Rust, or C++, Strong communication skills 📃 Skills: Python, Jax, Rust, C++, RL, LLM 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: The RL infrastructure team is looking for an engineer to help develop our RL training framework. RESPONSIBILITIES: - Design and implement the systems backing all RL workloads, from small scale ablations to production training runs - Profile, debug, and optimize end-to-end training performance - Improve scalability and observability of the RL stack BASIC QUALIFICATIONS: - Experience building, debugging, and optimizing efficiency of large-scale distributed systems - Comfortable diving into unfamiliar areas and solving problems at all levels of the stack - Proficiency in Python, Jax, Rust, and/or C++ PREFERRED SKILLS AND EXPERIENCE: - Experience with large scale LLM training infrastructure - Strong knowledge of reinforcement learning techniques - Experience with RL numerics COMPENSATION AND BENEFITS: - $180,000 - $440,000 USD - Equity package - Medical, vision, and dental coverage - 401(k) retirement plan - Short and long-term disability insurance - Life insurance - Additional discounts and perks SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.