June 30, 2026

Sr. Software Engineer (Data Center Automation)

Senior • On-site

Palo Alto, CA

About xAI

xAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity.

We operate with a flat organizational structure. All employees are expected to be hands-on and contribute directly to the company's mission. Strong work ethic, prioritization, and clear communication are essential.

About the Role

We are seeking a highly skilled Sr. Software Engineer to manage and enhance reliability across a multi-data center environment. This role focuses on automating processes, building robust observability solutions, and ensuring seamless operations for mission-critical AI infrastructure.

This position bridges software engineering principles with physical data center realities, prioritizing automation and observability to reduce downtime and improve recovery times. The team mitigates impact from scheduled and unscheduled maintenance through proactive automation and integrated reliability strategies.

Responsibilities

  • Design, develop, and deploy scalable services (primarily in Python and Rust) to automate monitoring, alerting, incident response, and infrastructure provisioning.
  • Implement and maintain observability practices including metrics, logging, tracing, and dashboards.
  • Collaborate with cross-functional teams to automate solutions for fault tolerance, disaster recovery, capacity planning, and environmental risk mitigation.
  • Troubleshoot complex issues across hardware, software, networking, and environmental systems while adhering to SLAs and error budgets.
  • Optimize Linux-based systems including kernel tuning, container orchestration, and automation scripting.
  • Analyze and troubleshoot large-scale, multi-data center network topologies.
  • Participate in on-call rotations, post-incident reviews, and continuous improvement initiatives.
  • Mentor junior engineers and document processes to strengthen automation culture.

Basic Qualifications

  • Bachelor's degree in a relevant technical field or equivalent experience.
  • 3+ years in SRE, Infrastructure, DevOps, or Systems Engineering in large-scale production environments.
  • Strong production programming experience in Python; familiarity with Rust or other systems languages.
  • Linux systems administration, performance tuning, and kernel-level knowledge.
  • Experience with Docker and Kubernetes.
  • Experience implementing monitoring, logging, tracing, and alerting systems.
  • Understanding of TCP/IP, routing, redundancy, and DNS.
  • Experience troubleshooting distributed systems and participating in incident response.
  • Ability to collaborate with cross-functional technical teams.

Preferred Skills and Experience

  • 5+ years in SRE or infrastructure roles in hyperscale or AI environments.
  • Experience scaling Kubernetes clusters with automation and high availability.
  • Proficiency in Rust.
  • Experience integrating software systems with physical data center infrastructure.
  • Experience building automated remediation and disaster recovery systems.
  • Optimization of Linux systems for AI workloads or GPU clusters.
  • Experience with bare-metal provisioning and multi-site failover mechanisms.
  • Mentoring and strong documentation skills.

xAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Similar jobs you might like

Technology

xAI

Sr. Software Engineer (Data Center Automation)

Senior

On-site

Memphis, TN

🏢 Summary: Senior Software Engineer role focused on improving reliability and automation across multi-data center AI infrastructure. The position combines strong software engineering skills with hands-on data center and systems expertise to build observability, automate remediation, and minimize downtime in mission-critical environments. The engineer will design scalable services, optimize Linux systems, and support distributed infrastructure with near-zero downtime requirements. 🗂️ Requirements: Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering or related field (or equivalent experience), 3+ years experience in SRE, Infrastructure, DevOps, or Systems Engineering in distributed production environments, Strong programming experience in Python, Solid Linux systems administration and performance tuning experience, Experience with Docker and Kubernetes or similar orchestration tools, Experience implementing observability solutions (metrics, logging, tracing, monitoring, alerting), Knowledge of networking fundamentals (TCP/IP, routing, DNS, redundancy), Experience troubleshooting distributed systems including hardware and network issues, Experience with on-call rotations and incident response practices, Ability to collaborate with cross-functional technical teams 📃 Skills: Python, Rust, Linux, Docker, Kubernetes, Prometheus, Grafana, TCP/IP, DNS, C++, Go 🏢 Description: ABOUT THE ROLE: We are seeking a highly skilled Sr. Software Engineer to join our team in managing and enhancing reliability across a multi-data center environment. This role focuses on automating processes, building and implementing robust observability solutions, and ensuring seamless operations for mission-critical AI infrastructure. The ideal candidate will combine strong coding abilities with hands-on data center experience to build scalable reliability services, optimize system performance, and minimize downtime—including close partnership with facility operations to address physical infrastructure impacts. In an era where AI workloads demand near-zero downtime, this position plays a pivotal role in bridging software engineering principles with physical data center realities. By prioritizing automation and observability, team members in this role can reduce mean time to recovery (MTTR) by up to 50% through proactive monitoring and automated remediation, based on industry benchmarks from high-scale environments like those at hyperscale cloud providers. The primary objective of this team is to mitigate downtime and minimize impact to end-users from both scheduled and unscheduled maintenance, as well as events affecting onsite data centers. This is achieved through proactive automation, robust observability, and integrated software-physical reliability strategies, ensuring our AI infrastructure remains resilient, scalable, and at the cutting edge of innovation. RESPONSIBILITIES: - Design, develop, and deploy scalable code and services (primarily in Python and Rust, with flexibility for emerging languages) to automate reliability workflows, including monitoring, alerting, incident response, and infrastructure provisioning. - Implement and maintain observability tools and practices, such as metrics collection, logging, tracing, and dashboards, to provide real-time insights into system health across multiple data centers. - Collaborate with cross-functional teams—including software development, network engineering, site operations, and facility operations—to identify reliability bottlenecks and automate solutions for fault tolerance, disaster recovery, capacity planning, and physical/environmental risk mitigation. - Troubleshoot and resolve complex issues in data center environments, including hardware failures, environmental anomalies, software bugs, and network-related problems, while adhering to reliability principles like error budgets and SLAs. - Optimize Linux-based systems for performance, security, and reliability, including kernel tuning, container orchestration (e.g., Kubernetes), and scripting for automation. - Understand network topologies and concepts in large-scale, multi-data center environments to troubleshoot connectivity, routing, redundancy, and performance issues. - Participate in on-call rotations, post-incident reviews (blameless postmortems), and continuous improvement initiatives to enhance overall site reliability. - Mentor junior team members and document processes to foster a culture of automation and knowledge sharing. BASIC QUALIFICATIONS: - Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or a closely related technical field (or equivalent professional experience). - 3+ years of hands-on experience in site reliability engineering (SRE), infrastructure engineering, DevOps, or systems engineering in large-scale, distributed, or production environments. - Strong programming skills with proven production experience in Python; experience with Rust or strong fundamentals in a systems-level language (e.g., Go, C++). - Solid experience with Linux systems administration, performance tuning, kernel-level understanding, and scripting/automation in production environments. - Practical knowledge of containerization and orchestration technologies, such as Docker and Kubernetes (or similar systems). - Experience implementing observability solutions, including metrics, logging, tracing, monitoring tools (e.g., Prometheus, Grafana), alerting, and dashboards. - Familiarity with troubleshooting complex issues in distributed systems, including software bugs, hardware failures, network problems, and environmental factors. - Understanding of networking fundamentals (TCP/IP, routing, redundancy, DNS) in large-scale or multi-site environments. - Experience participating in on-call rotations, incident response, post-incident reviews (blameless postmortems), and reliability practices such as error budgets or SLAs. - Ability to collaborate effectively with cross-functional technical teams. PREFERRED SKILLS AND EXPERIENCE: - 5+ years of experience in SRE or infrastructure roles in hyperscale, cloud, or AI/ML training infrastructure environments. - Hands-on experience operating or scaling Kubernetes clusters at large scale. - Proficiency in Rust for systems programming and performance-critical components. - Experience integrating software reliability tools with physical data center infrastructure (power, cooling, environmental monitoring). - Experience building automated remediation, fault tolerance, disaster recovery, capacity planning, or predictive failure detection systems. - Background in optimizing Linux-based systems for AI workloads, GPU clusters, or high-throughput compute environments. - Experience with bare-metal provisioning, data center interconnects, or hybrid/multi-site failover mechanisms. - Mentoring experience and strong documentation skills.

Technology

xAI

Backend Engineer - API

Senior

On-site

Palo Alto, CA

🏢 Summary: Engineering role focused on building and owning a high-throughput, low-latency API and backend infrastructure for large-scale model inference. The position involves designing and operating reliable, horizontally scalable distributed systems that serve billions of tokens per minute. You will develop and maintain model serving, routing, SDKs, and observability within a production-grade environment. 🗂️ Requirements: Expert knowledge of Rust or C++, Experience building and maintaining horizontally scalable distributed systems, Experience designing reliable high-availability production infrastructure, Knowledge of observability and reliability best practices, Experience operating PostgreSQL, Clickhouse, or MongoDB 📃 Skills: Rust, C++, Go, PostgreSQL, Clickhouse, MongoDB, gRPC, Docker, Kubernetes, TensorRT, vLLM, SGLang, REST, SDK 🏢 Description: About xAI xAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.ABOUT THE ROLE: As an ideal candidate you have a good understanding of how highly scalable and reliable production infrastructure is built. Most of our backend infrastructure is written in Rust. So familiarity with a compiled language such as C++, Rust, or Go is highly beneficial. RESPONSIBILITIES: Build the xAI API that serves our models to developers worldwide Own the end-to-end system responsible for high-throughput inference, handling billions of tokens per minute with low latency and high availability, including model serving infrastructure, request routing, SDK development, rate limiting, observability, and efficient scaling BASIC QUALIFICATIONS: Expert knowledge of either Rust or C++ Experience in designing, implementing, and maintaining reliable and horizontally scalable distributed systems Knowledge of service observability and reliability best practices Experience in operating commonly used databases such as PostgreSQL, Clickhouse, and MongoDB PREFERRED SKILLS AND EXPERIENCE: Experience with LLM inference engines and serving frameworks (e.g., SGLang, TensorRT, vLLM) Experience designing or building with agent SDKs and agent orchestration frameworks Experience with Docker, Kubernetes, and containerized applications Expert knowledge of gRPC (unary, response streaming, bi-directional streaming, REST mapping) COMPENSATION AND BENEFITS $180,000 - $440,000 USD Base salary is just one part of our total rewards package at xAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.xAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Technology

xAI

Software Engineer - Kernels/CUDA (C++)

Senior

On-site

Seattle, WA

🏢 Summary: High-impact compute infrastructure role focused on building and optimizing massive GPU supercomputers for AI training and inference. The position involves low-level systems programming, GPU kernel optimization, Linux kernel internals, orchestration, and distributed infrastructure to improve scalability, reliability, and performance of AI workloads. 🗂️ Requirements: Deep systems programming experience in C/C++ or Rust, Experience with large-scale GPU clusters or distributed compute infrastructure, Hands-on GPU kernel optimization using CUTLASS, custom kernels, or Nsight, Knowledge of Linux kernel internals, scheduling, virtualization, or orchestration, Experience building high-performance AI training or inference infrastructure, Ability to optimize memory-bound and compute-bound workloads, Experience operating exabyte-scale storage systems 📃 Skills: CUDA, CUTLASS, Tensor, Nsight, Linux, KVM, Firecracker, Kubernetes, C++, Rust 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.ABOUT THE ROLE: We are building one of the world's largest AI supercomputers from the ground up. As part of the Compute Infrastructure team, you will own both the raw GPU supercomputer and the platform layer that runs on top of it. You will work across the full stack — from low-level GPU kernel optimizations and Linux kernel internals to massive-scale orchestration and virtualization — to make training and inference at SpaceXAI as fast, reliable, and scalable as possible. This is a broad, high-impact role that combines hardcore supercompute and compute infrastructure work. Your contributions will directly accelerate Grok's training speed and overall AI progress. RESPONSIBILITIES: Design, build, and optimize massive GPU clusters for extreme-scale training and inference workloads Develop and tune low-level CUDA kernels (GeMM, Attention, etc.), using CUTLASS, Tensor Cores, and Nsight for maximum performance Profile, debug, and eliminate bottlenecks across GPU memory hierarchy, networking fabric, filesystems, and multi-GPU operation Collaborate closely with AI research teams to deliver production-grade performance and scalability PREFERRED SKILLS AND EXPERIENCE: Deep low-level systems programming (C/C++/PTX/SASS) Strong experience with large-scale GPU clusters or distributed compute infrastructure at production scale Hands-on work with GPU kernel optimization (CUTLASS, custom kernels, Nsight profiling) Track record of building or running high-performance infrastructure for AI workloads (training or inference platforms) Ability to reason from first principles and optimize for both memory-bound and compute-bound scenarios COMPENSATION AND BENEFITS: $180,000 - $440,000 USD Base salary is just one part of our total rewards package at SpaceXAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Technology

SpaceXAI

Software Engineer - Data

Mid

On-site

Palo Alto, CA

150,000 - 249,996 USD/yr

🏢 Summary: Software Engineer role focused on building scalable data platforms and pipelines for AI model training, including data acquisition, preparation, orchestration, monitoring, and quality evaluation. The position involves collaborating with ML and data teams to develop reliable systems that support large-scale training workflows and improve model performance. 🗂️ Requirements: Bachelor's degree in Computer Science, Data Science, Engineering, Math, Physics, or related field OR 2+ years of professional software experience, 1+ years of experience in software engineering, application development, data engineering, or data science, Experience with programming languages such as Python, Rust, Java, C#, Scala, or Go, Experience with frontend frameworks such as Angular or React, Hands-on experience with Kubernetes and containerized deployments, Experience with relational and non-relational databases, Understanding of version control, testing, CI/CD, deployment, and monitoring, Understanding of statistics and machine learning frameworks 📃 Skills: Python, Rust, Java, C#, Scala, Go, Angular, React, Kubernetes, Ray, PostgreSQL, Iceberg, Clickhouse, Grafana, Superset, CI/CD, Docker, SQL, MachineLearning 🏢 Description: ABOUT THE ROLE: At xAI, we are building AI systems that push the frontier of human knowledge and scientific discovery. High-quality data is fundamental to every stage of that mission. Our Data team is responsible for ensuring that the models are trained on the right data, in the right form, at the right quality, across every phase of the training lifecycle. This includes partnering closely with acquisition teams to identify where valuable data can be sourced, determining what data is needed to improve model performance, and building the production pipelines and systems that transform raw inputs into high-quality training data at scale. We work at the intersection of software, data, infrastructure, and machine learning to ensure our models train effectively and reliably. As a Software Engineer on xAI's Data team, you will be responsible for developing applications that power data acquisition, preparation, training, quality evaluation, and delivery for model training. You will provide the ability to run training in a reliable, scalable and repeatable manner. You will also provide visibility on training status and data lineage. You will work closely with acquisition teams, ML engineers, and data engineers to build a reliable data pipeline to run training at scale. The ideal candidate combines strong software engineering fundamentals and excellent coding practices. RESPONSIBILITIES: - Develop a highly reliable and scalable enterprise data platform to orchestrate data acquisition, preparation, training, quality evaluation, and delivery for model training - Create new features such as data lineage, visibility, and monitoring for end-to-end training that improve the quality of the data and model performance - Collaborate with peers on architecture, design, and code reviews - Build prototypes to prove out key design concepts and quantify technical constraints - Own all aspects of software engineering and product development - Deep dive into business problems, find efficient solutions and apply first principles thinking BASIC QUALIFICATIONS: - Bachelor's degree in computer science, data science, engineering, math, physics, or scientific discipline; OR 2+ years of professional experience building software in lieu of a degree - 1+ years of experience in application development, software engineering, data engineering, or data science PREFERRED SKILLS AND EXPERIENCE: - Programming experience in Python, Rust, Java, C#, Scala, Go or similar languages - Frontend experience in Angular, React, or similar JavaScript frameworks - Hands-on experience with Kubernetes and containerized deployments - Experience with Ray, AI training and orchestration - Experience with relational and non-relational databases, data lakes e.g. PostgreSQL, Iceberg, Clickhouse, or similar - Experience with data exploration tools like Grafana, Superset, or similar - Good understanding of version control, testing, continuous integration, build, deployment and monitoring - Good understanding of statistics, machine learning algorithms and frameworks COMPENSATION AND BENEFITS: - $150,000 - $250,000 USD - Base salary is part of a total rewards package that includes equity, medical, vision, and dental coverage, 401(k) retirement plan, disability insurance, life insurance, and additional perks.

Technology

xAI

Site Reliability Engineer - Cybersecurity

Senior

On-site

Palo Alto, CA

🏢 Summary: Cybersecurity / SRE role focused on securing and maintaining the reliability of a large-scale fintech platform operating in hybrid cloud environments. The position emphasizes Kubernetes and container security, SIEM management, CI/CD protection, and automation using Python and infrastructure-as-code tools. Candidates will work on mission-critical distributed systems, ensuring regulatory compliance and resilient security operations at scale. 🗂️ Requirements: Experience securing hybrid AWS/on-premises environments, Strong proficiency in Python, Strong proficiency in Terraform, Strong proficiency in Puppet, Deep expertise in Kubernetes, Experience with container security, Hands-on experience with GitHub Actions, Experience with Prometheus, Experience with Grafana, Experience with CloudWatch, Experience with Karma, Experience managing and integrating Wazuh, Experience with security scanning tools (Semgrep, Trivy, Falco), Experience with IAM and security posture management, Ability to comply with PCI and NIST CSF standards, Located in SF Bay Area or willing to relocate 📃 Skills: AWS, IAM, Python, Terraform, Puppet, Kubernetes, Docker, GitHub, Prometheus, Grafana, CloudWatch, Karma, Wazuh, Semgrep, Trivy, Falco, PCI, NIST, CI/CD 🏢 Description: ABOUT xAI xAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: The Cybersecurity / SRE team is focused on ensuring the security and reliability of X Money. This role will primarily focus on the X Money platform but will also cross over with the X Social platform. The ideal candidate will have experience in the banking, money transmission, and P2P payments industry. We emphasize working with large distributed systems and security platforms at scale, with an automation-first mindset. You'll be responsible for securing and maintaining the reliability of X Money's infrastructure. You'll work closely with cross-functional teams to enhance security measures, improve system resilience, and implement best practices. RESPONSIBILITIES: - Build and secure mission-critical applications in a hybrid cloud environment. - Manage identities and roles effectively. - Monitor and remediate infrastructure to comply with regulations and best practices (e.g., PCI, NIST CSF). - Maintain a SIEM and all data pipelines needed for reliable alerting. - Design and implement secure container standards and automation to enable frictionless developer workflows. - Maintain Kubernetes security aligned with current best practices. - Build, deploy, and maintain security operations infrastructure using Python, Terraform, and Puppet. - Secure and enhance CI/CD pipelines. - Integrate and maintain code scanning platforms. - Develop dashboards and alerts from security metrics. - Own security projects: identify issues and implement solutions. - Apply critical analysis and problem-solving skills. BASIC QUALIFICATIONS: - Proven experience securing hybrid AWS/on-premises environments, including IAM and overall security posture. - Strong proficiency in Python, Terraform, and Puppet. - Certifications like CISA, CRISC, CGEIT, Security+, CASP+, or similar preferred. - Deep expertise in Kubernetes and container security. - Hands-on expertise building GitHub Actions and workflows. - Extensive experience with Prometheus, Grafana, CloudWatch, and Karma. - Well versed in management and integrations of Wazuh. - Hands-on experience with security scanning tools (Semgrep, Trivy, Falco). - Proactive mindset with strong ownership and problem-solving skills. - Excellent critical thinking and analytical abilities. - Located in the SF Bay Area or willing to relocate. COMPENSATION AND BENEFITS: $180,000 - $440,000 USD Base salary is just one part of our total rewards package, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks. xAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Technology

xAI

Software Engineer - Networking Software and Services

Senior

On-site

Palo Alto, CA

🏢 Summary: Opportunity to build and scale automation-first network software and services supporting large-scale GPU supercomputing fabrics for AI training and inference. The role focuses on developing tools for network management, metrics collection, provisioning, monitoring, and auto-remediation while implementing Infrastructure as Code best practices. You will design highly scalable, reliable systems that orchestrate tens of thousands of network devices in production environments. 🗂️ Requirements: Deep experience working with network engineers and network topologies, Strong knowledge of physical and logical network architectures, Strong knowledge of network protocols, Proven experience designing scalable and reliable software systems, Experience building systems that orchestrate large-scale network devices, Ability to implement Infrastructure as Code best practices, Experience enhancing deployment pipelines, Ability to create and define meaningful metrics for prioritization, Strong communication skills 📃 Skills: Python, Go, TCP/IP, BGP, RDMA, IaC, Networking, Automation 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.ABOUT THE ROLE: As part of the Network Software and Services for AI (nssAI) team at SpaceXAI, you'll build cutting-edge software, services, and frameworks to empower our Network Development Engineers. Working hands-on, you'll tackle all facets of network management—metric collection, configuration, zero-touch provisioning, monitoring, and auto-remediation—driving automation-first solutions for SpaceXAI's production and ancillary networks. Expect to develop extensible tools, streamline complex processes, and ensure rock-solid reliability to support SpaceXAI's mission of accelerating human scientific discovery through AI. RESPONSIBILITIES: Building software and tools with extensive metrics coverage for some of the world's largest GPU supercomputing network fabrics used for AI training and serving customer inference queries. Implement IaC best practices, enhancing deployment pipelines, and ensuring robust, secure service delivery across our production environments. BASIC QUALIFICATIONS: Deep experience collaborating with network engineers daily using extensive knowledge of network topologies, physical and logical, and network protocols. Expert knowledge and proven history with designing scalable and reliable software from the ground up that can build and orchestrate tens of thousands of network devices at lightning speeds. Ability to thrive in ambiguity, creating metrics that will help prioritize the focus of the team and your own. COMPENSATION AND BENEFITS: $150,000 - 250,000k Base Base salary is just one part of our total rewards package at SpaceXAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Technology

SpaceXAI

Data Engineer

Junior

On-site

Palo Alto, CA

150,000 - 210,000 USD/yr

🏢 Summary: Data Engineer / AI Engineer role focused on building scalable data pipelines, improving training data quality, and supporting AI model development through statistical analysis and production-grade systems. The position involves collaborating with ML and software teams to optimize data acquisition, processing, and evaluation for large-scale neural network training. 🗂️ Requirements: Bachelor's degree in Computer Science, Data Science, Physics, Mathematics, or STEM field, 1+ years of data or software engineering experience, Experience implementing or analyzing language models or neural networks, Strong Python programming skills, Experience with production data pipelines, Knowledge of machine learning and statistical analysis, Familiarity with large-scale datasets and distributed systems, Ability to work in fast-paced, evolving technical environments 📃 Skills: Python, Kubernetes, Parquet, MachineLearning, NeuralNetworks, Analytics, DataEngineering, Statistics, Clustering, Forecasting, AnomalyDetection 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: At xAI, we are building AI systems that push the frontier of human knowledge and scientific discovery. High-quality data is fundamental to every stage of that mission. Our Data team is responsible for ensuring that the models are trained on the right data, in the right form, at the right quality, across every phase of the training lifecycle. This includes partnering closely with acquisition teams to identify where valuable data can be sourced, determining what data is needed to improve model performance, and building the production pipelines and systems that transform raw inputs into high-quality training data at scale. We work at the intersection of data, infrastructure, and machine learning to ensure our models train effectively and reliably. As a Data Engineer / AI Engineer on xAI's Data team, you will be responsible for developing the systems, processes, and production code that power data acquisition, preparation, quality evaluation, and delivery for model training. You will work closely with acquisition teams, ML engineers, and software engineers to identify data needs, build scalable data pipelines, and continuously improve the quality of the data that shapes model behavior. The ideal candidate combines strong software engineering fundamentals and excellent coding practices with deep intuition for statistics, neural networks, and how data quality influences training outcomes. RESPONSIBILITIES: - Analyze the performance and impact of data used throughout the model training lifecycle - Investigate anomalous model behavior and rigorously identify the data issues that drive poor downstream performance - Design, build, and improve the data cleaning, transformation, and quality-control steps required to produce high-quality training data - Research, evaluate, and develop frontier methods for improving data quality and effectiveness in AI model development - Apply statistical techniques and empirical analysis to make informed, data-driven decisions about dataset quality and model outcomes - Partner across teams to identify where data needs exist and define the highest-impact opportunities for new data acquisition and improvement - Build and maintain production-grade data pipelines, tooling, and software systems that ingest, process, validate, and deliver data for training - Develop metrics, evaluation frameworks, and monitoring systems to assess how data quality influences model behavior at scale - Fuse data from multiple sources into reliable, usable datasets for research and production model training - Create shared datasets, tooling, and internal data products that enable other teams to analyze, debug, and improve model performance BASIC QUALIFICATIONS: - Bachelor's degree in computer science, data science, physics, mathematics, or a STEM discipline - 1+ years of data/software engineering experience (internship experience is applicable) - Experience in implementing or analyzing language models or neural networks PREFERRED SKILLS AND EXPERIENCE: - Professional experience in analytics, data science, machine learning, or data engineering - Experience building and operating production data pipelines for neural network or large-scale machine learning workloads - Strong experience with Python and the broader ecosystem of libraries and tools used in modern machine learning and data development - Experience working with Parquet or similar columnar storage formats in large-scale data systems - Familiarity with Kubernetes and distributed production environments - Experience developing predictive models and machine learning pipelines, including clustering, forecasting, anomaly detection, or related techniques - Experience working with very large-scale datasets, including terabyte- to petabyte-scale data systems - Strong statistical intuition and the ability to use quantitative analysis to guide technical and product decisions - Ability to operate effectively in a dynamic environment with evolving priorities, changing requirements, and fast-moving technical challenges - Demonstrated ability to take ownership of ambiguous problems, drive projects independently, and develop new expertise where needed COMPENSATION AND BENEFITS: - $150,000 - $210,000 USD - Equity package - Medical, vision, and dental coverage - 401(k) retirement plan - Short and long-term disability insurance - Life insurance - Additional discounts and perks SpaceXAI is an equal opportunity employer. For details on data processing, view the Recruitment Privacy Notice.

Technology

xAI

Senior Data Analyst- Fraud & AML

Senior

On-site

Palo Alto, CA

12,333 - 18,333 USD/yr

🏢 Summary: Senior Data Scientist role focused on designing and optimizing AML and fraud detection models to strengthen financial crime compliance and transaction monitoring. The position involves building advanced analytics solutions, coverage assessment frameworks, and automated reporting to support BSA/AML, OFAC, and regulatory requirements. It is a cross-functional, high-impact role combining machine learning, compliance expertise, and scalable data solutions. 🗂️ Requirements: 7+ years data science experience in financial services, 4+ years experience in fraud and financial crime compliance, Master's degree in quantitative field, Experience building transaction monitoring models in regulated environment, Strong knowledge of BSA/AML and SAR processes, Understanding of sanctions screening and model risk management, Proficiency in Python and SQL, Experience supporting regulatory examinations, U.S. work authorization under ITAR requirements 📃 Skills: Python, SQL, MachineLearning, Statistics, AML, BSA, OFAC, SAR, FraudDetection, TransactionMonitoring, Compliance, ModelRisk, DataAnalytics, Dashboards, Automation, RPA 🏢 Description: About xAI xAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.ABOUT THE ROLE: We are looking for a Senior Data Scientist to join our Compliance Program and play a pivotal role in modernizing and strengthening our financial crime detection capabilities. You will architect, build, and optimize data-driven transaction monitoring models, coverage assessment frameworks, and advanced analytics solutions that directly support BSA/AML regulatory compliance, including but not limited to SAR filing, Customer Identification Program elements, and Enhanced Due Diligence measures. The role will also support the OFAC Sanctions, Fraud and overall risk prioritization across multiple products and jurisdictions. This is a high-impact, cross-functional role that blends advanced analytics, machine learning, and deep compliance domain expertise. You will work closely with Compliance, Engineering, Model Risk, Product, and external regulators to ensure our controls are robust, defensible, and scalable. RESPONSIBILITIES: Design, develop, and enhance AML and fraud models, rules, and heuristics using Python, SQL, and AI-enabled tooling; partner with the Compliance Machine Learning team on model reviews to improve detection rates and reduce false positives. Build and maintain interactive performance dashboards and automated reporting solutions that track key risk, productivity, and capacity metrics for senior leadership and regulators. Architect and implement enterprise-wide Transaction Monitoring Coverage Assessment frameworks, including standardized methodologies for gap identification, root-cause analysis, remediation planning, and ongoing sustainability monitoring. Lead complex data initiatives, including extraction of SAR filing metrics with product-level breakdowns and development of jurisdiction- and typology-specific SAR narrative generator tools. Embed data science best practices into product launches and feature rollouts to proactively identify and close monitoring coverage gaps. Support regulatory examinations (e.g., NYDFS Part 504) by preparing analytical documentation, third-party validation materials, and executive certification packages. Drive continuous improvement of compliance operations through automation, process optimization, and advanced analytics. BASIC QUALIFICATIONS: 7+ years of hands-on data science / advanced analytics experience in financial services, with at least 4 years focused on fraud and financial crime compliance. Master's degree (or higher) in Applied Mathematics, Statistics, Data Science, Actuarial Science, or a related quantitative field. Proven track record of building and optimizing transaction monitoring models, coverage frameworks, or compliance analytics programs in a regulated environment (fintech, bank, or payment company preferred). Deep understanding of BSA/AML regulations, suspicious activity reporting, customer due diligence, sanctions screening, and model risk management principles. Demonstrated ability to translate complex regulatory requirements into actionable data solutions and present findings to senior leadership and regulators. Certified Anti-Money Laundering Specialist (CAMS) or equivalent compliance certification is strongly preferred. PREFERRED SKILLS AND EXPERIENCE: Experience leading cross-functional initiatives involving Engineering, Legal, Product Compliance, and external consulting partners. Background in building internal case management systems, SAR automation tools, or RPA solutions. Familiarity with AML detection platforms. Track record of delivering measurable impact (e.g., reduced case volumes, improved detection of high-risk activity, increased operational efficiency). COMPENSATION AND BENEFITS: $148,000- $220,000 USD Base salary is just one part of our total rewards package at xAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks. ITAR REQUIREMENTS: To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State. Learn more about the ITAR here. xAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Technology

New offer

xAI

Software Engineer– X Core Product

Mid

On-site

Palo Alto, CA

🏢 Summary: Software Engineer role focused on building and scaling a high-volume platform serving 600M+ users, with ownership across backend services, infrastructure, AI integrations, and fullstack features. The position involves designing low-latency distributed systems, collaborating across mobile and web teams, and driving architecture, scalability, and reliability decisions. 🗂️ Requirements: 2+ years of experience with large-scale consumer applications, Proficiency in distributed systems, Experience with high-scale, low-latency environments, Knowledge of Rust, Knowledge of Go, Knowledge of Python, Knowledge of Java, Experience with high-volume streaming systems, Ability to build backend services and APIs, Experience with microservices architecture 📃 Skills: Rust, Go, Python, Java, APIs, Microservices, DistributedSystems, Streaming, iOS, Android, Web, Backend, Fullstack, Analytics, AI 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: As a Software Engineer for X Platform, you'll join the thirty person team responsible for building and scaling X. You will be tasked with independently owning significant parts of the system end-to-end: from intuitive user interfaces to robust backend services, data infrastructure, and deep AI integrations. This is a unique opportunity to impact while creating products that are loved by 600M+ users globally. RESPONSIBILITIES: - Develop backend services, APIs, and data models to support high-volume, multi-user environments. - Work with iOS, Android & Web client engineers to ship products. - Design robust infrastructure and microservices for payments, transactions, growth, monetization, and engagement across platforms. - Build and maintain fullstack features, including user dashboards, personalized experiences, content delivery, interactive tools, assessments, and real-time analytics. - Lead architecture, scalability, and reliability decisions for high-concurrency, low-latency systems. - Uphold engineering excellence via testing, monitoring, deployment, and secure data handling. BASIC QUALIFICATIONS: - Proficiency in distributed systems for high-scale, low-latency environments; languages like Rust, Go, Python & Java, and high volume streaming systems. - 2+ years of experience working on large scale consumer applications. PREFERRED SKILLS AND EXPERIENCE: - 5+ years of experience working on large scale consumer applications or early-mid stage startup experience as a founding engineer, emphasizing rapid prototyping, user-centric design, and AI solutions. COMPENSATION AND BENEFITS: - $180,000 - $440,000 USD - Base salary is just one part of the total rewards package, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks. SpaceXAI is an equal opportunity employer. For details on data processing, view the Recruitment Privacy Notice.

Technology

SpaceXAI

Software Engineer– X Core Product

Mid

On-site

Palo Alto, CA

129,996 - 234,996 USD/yr

🏢 Summary: Software Engineer role focused on building and scaling a high-volume consumer platform with end-to-end ownership across frontend, backend, infrastructure, and AI integrations. The position involves developing distributed systems, microservices, analytics, and real-time features for products used by hundreds of millions of users. Candidates should have experience with large-scale applications, low-latency systems, and modern backend technologies. 🗂️ Requirements: 2+ years experience with large-scale consumer applications, Proficiency in distributed systems, Experience with high-scale low-latency environments, Experience with backend services and APIs, Experience with high-volume streaming systems, Strong communication skills, Ability to work in high-concurrency environments 📃 Skills: Rust, Go, Python, Java, APIs, Microservices, iOS, Android, Web, AI, Streaming, Analytics 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: As a Software Engineer for X Product/Platform, you'll join the thirty person team responsible for building and scaling X. You will be tasked with independently owning significant parts of the system end-to-end: from intuitive user interfaces to robust backend services, data infrastructure, and deep AI integrations. This is a unique opportunity to impact while creating products that are loved by 600M+ users globally. RESPONSIBILITIES: - Develop backend services, APIs, and data models to support high-volume, multi-user environments. - Work with iOS, Android & Web client engineers to ship products. - Design robust infrastructure and microservices for payments, transactions, growth, monetization, and engagement across platforms. - Build and maintain fullstack features, including user dashboards, personalized experiences, content delivery, interactive tools, assessments, and real-time analytics. - Lead architecture, scalability, and reliability decisions for high-concurrency, low-latency systems. - Uphold engineering excellence via testing, monitoring, deployment, and secure data handling. BASIC QUALIFICATIONS: - Proficiency in distributed systems for high-scale, low-latency environments; languages like Rust, Go, Python & Java, and high volume streaming systems. - 2+ years of experience working on large scale consumer applications. PREFERRED SKILLS AND EXPERIENCE: - 5+ years of experience working on large scale consumer applications or early-mid stage startup experience as a founding engineer, emphasizing rapid prototyping, user-centric design, and AI solutions. COMPENSATION AND BENEFITS: - $180,000 - $440,000 USD - Base salary is just one part of the total rewards package, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks. SpaceXAI is an equal opportunity employer. For details on data processing, view the Recruitment Privacy Notice.