July 3, 2026
Software Engineer - Kernels/CUDA (C++)
Senior • On-site
Palo Alto, CA
SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.
ABOUT THE ROLE:
We are building one of the world's largest AI supercomputers from the ground up. As part of the Compute Infrastructure team, you will own both the raw GPU supercomputer and the platform layer that runs on top of it. You will work across the full stack — from low-level GPU kernel optimizations and Linux kernel internals to massive-scale orchestration and virtualization — to make training and inference at SpaceXAI as fast, reliable, and scalable as possible.
This is a broad, high-impact role that combines hardcore supercompute and compute infrastructure work. Your contributions will directly accelerate Grok's training speed and overall AI progress.
RESPONSIBILITIES:
- Design, build, and optimize massive GPU clusters for extreme-scale training and inference workloads
- Develop and tune low-level CUDA kernels (GeMM, Attention, etc.), using CUTLASS, Tensor Cores, and Nsight for maximum performance
- Profile, debug, and eliminate bottlenecks across GPU memory hierarchy, networking fabric, filesystems, and multi-GPU operation
- Collaborate closely with AI research teams to deliver production-grade performance and scalability
PREFERRED SKILLS AND EXPERIENCE:
- Deep low-level systems programming (C/C++/PTX/SASS)
- Strong experience with large-scale GPU clusters or distributed compute infrastructure at production scale
- Hands-on work with GPU kernel optimization (CUTLASS, custom kernels, Nsight profiling)
- Track record of building or running high-performance infrastructure for AI workloads (training or inference platforms)
- Ability to reason from first principles and optimize for both memory-bound and compute-bound scenarios
COMPENSATION AND BENEFITS:
$180,000 - $440,000 USD
Base salary is just one part of our total rewards package at SpaceXAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.
SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.
Similar jobs you might like
Technology

xAI
Software Engineer - Kernels/CUDA (C++)
Senior
On-site
Seattle, WA
🏢 Summary: High-impact compute infrastructure role focused on building and optimizing massive GPU supercomputers for AI training and inference. The position involves low-level systems programming, GPU kernel optimization, Linux kernel internals, orchestration, and distributed infrastructure to improve scalability, reliability, and performance of AI workloads. 🗂️ Requirements: Deep systems programming experience in C/C++ or Rust, Experience with large-scale GPU clusters or distributed compute infrastructure, Hands-on GPU kernel optimization using CUTLASS, custom kernels, or Nsight, Knowledge of Linux kernel internals, scheduling, virtualization, or orchestration, Experience building high-performance AI training or inference infrastructure, Ability to optimize memory-bound and compute-bound workloads, Experience operating exabyte-scale storage systems 📃 Skills: CUDA, CUTLASS, Tensor, Nsight, Linux, KVM, Firecracker, Kubernetes, C++, Rust 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.ABOUT THE ROLE: We are building one of the world's largest AI supercomputers from the ground up. As part of the Compute Infrastructure team, you will own both the raw GPU supercomputer and the platform layer that runs on top of it. You will work across the full stack — from low-level GPU kernel optimizations and Linux kernel internals to massive-scale orchestration and virtualization — to make training and inference at SpaceXAI as fast, reliable, and scalable as possible. This is a broad, high-impact role that combines hardcore supercompute and compute infrastructure work. Your contributions will directly accelerate Grok's training speed and overall AI progress. RESPONSIBILITIES: Design, build, and optimize massive GPU clusters for extreme-scale training and inference workloads Develop and tune low-level CUDA kernels (GeMM, Attention, etc.), using CUTLASS, Tensor Cores, and Nsight for maximum performance Profile, debug, and eliminate bottlenecks across GPU memory hierarchy, networking fabric, filesystems, and multi-GPU operation Collaborate closely with AI research teams to deliver production-grade performance and scalability PREFERRED SKILLS AND EXPERIENCE: Deep low-level systems programming (C/C++/PTX/SASS) Strong experience with large-scale GPU clusters or distributed compute infrastructure at production scale Hands-on work with GPU kernel optimization (CUTLASS, custom kernels, Nsight profiling) Track record of building or running high-performance infrastructure for AI workloads (training or inference platforms) Ability to reason from first principles and optimize for both memory-bound and compute-bound scenarios COMPENSATION AND BENEFITS: $180,000 - $440,000 USD Base salary is just one part of our total rewards package at SpaceXAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.
Technology

SpaceXAI
Software Engineer - Kernels/CUDA (C++)
Senior
On-site
Palo Alto, CA
129,996 - 234,996 USD/yr
🏢 Summary: Role focused on building and optimizing large-scale GPU supercomputing infrastructure for AI training and inference. The position involves low-level GPU kernel optimization, distributed compute systems, and full-stack performance tuning across hardware and platform layers. Candidates will work closely with AI research teams to improve scalability, reliability, and training speed. 🗂️ Requirements: Deep low-level systems programming experience, Experience with large-scale GPU clusters, Hands-on GPU kernel optimization, Experience with distributed compute infrastructure, Ability to optimize memory-bound and compute-bound workloads, Strong communication skills 📃 Skills: CUDA, C, C++, PTX, SASS, CUTLASS, Nsight, Linux, GPU, Tensor 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: We are building one of the world's largest AI supercomputers from the ground up. As part of the Compute Infrastructure team, you will own both the raw GPU supercomputer and the platform layer that runs on top of it. You will work across the full stack — from low-level GPU kernel optimizations and Linux kernel internals to massive-scale orchestration and virtualization — to make training and inference as fast, reliable, and scalable as possible. This is a broad, high-impact role that combines hardcore supercompute and compute infrastructure work. Your contributions will directly accelerate Grok's training speed and overall AI progress. RESPONSIBILITIES: - Design, build, and optimize massive GPU clusters for extreme-scale training and inference workloads - Develop and tune low-level CUDA kernels (GeMM, Attention, etc.), using CUTLASS, Tensor Cores, and Nsight for maximum performance - Profile, debug, and eliminate bottlenecks across GPU memory hierarchy, networking fabric, filesystems, and multi-GPU operation - Collaborate closely with AI research teams to deliver production-grade performance and scalability PREFERRED SKILLS AND EXPERIENCE: - Deep low-level systems programming (C/C++/PTX/SASS) - Strong experience with large-scale GPU clusters or distributed compute infrastructure at production scale - Hands-on work with GPU kernel optimization (CUTLASS, custom kernels, Nsight profiling) - Track record of building or running high-performance infrastructure for AI workloads (training or inference platforms) - Ability to reason from first principles and optimize for both memory-bound and compute-bound scenarios COMPENSATION AND BENEFITS: - $180,000 - $440,000 USD - Equity - Medical, vision, and dental coverage - 401(k) retirement plan - Short and long-term disability insurance - Life insurance - Additional discounts and perks SpaceXAI is an equal opportunity employer. For details on data processing, view the Recruitment Privacy Notice.
Technology

SpaceXAI
Software Engineer - Kernels/CUDA (C++)
Senior
On-site
Seattle, WA
123,996 - 170,004 USD/yr
🏢 Summary: Role focused on building and optimizing large-scale AI supercomputing infrastructure for training and inference workloads. The position involves low-level GPU kernel optimization, distributed systems engineering, and full-stack compute infrastructure development. Candidates will work on performance, scalability, and reliability of massive GPU clusters supporting AI model training. 🗂️ Requirements: Deep systems programming experience in C/C++/PTX/SASS, Experience with large-scale GPU clusters or distributed compute infrastructure, Hands-on GPU kernel optimization experience, Experience with CUTLASS and Nsight profiling, Track record building high-performance AI infrastructure, Ability to optimize memory-bound and compute-bound workloads, Strong communication skills, Ability to work hands-on in a high-impact engineering environment 📃 Skills: CUDA, C, C++, PTX, SASS, CUTLASS, Nsight, Linux, GPU, Tensor, GeMM, Attention 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: We are building one of the world's largest AI supercomputers from the ground up. As part of the Compute Infrastructure team, you will own both the raw GPU supercomputer and the platform layer that runs on top of it. You will work across the full stack — from low-level GPU kernel optimizations and Linux kernel internals to massive-scale orchestration and virtualization — to make training and inference as fast, reliable, and scalable as possible. This is a broad, high-impact role that combines hardcore supercompute and compute infrastructure work. Your contributions will directly accelerate Grok's training speed and overall AI progress. RESPONSIBILITIES: - Design, build, and optimize massive GPU clusters for extreme-scale training and inference workloads - Develop and tune low-level CUDA kernels (GeMM, Attention, etc.), using CUTLASS, Tensor Cores, and Nsight for maximum performance - Profile, debug, and eliminate bottlenecks across GPU memory hierarchy, networking fabric, filesystems, and multi-GPU operation - Collaborate closely with AI research teams to deliver production-grade performance and scalability PREFERRED SKILLS AND EXPERIENCE: - Deep low-level systems programming (C/C++/PTX/SASS) - Strong experience with large-scale GPU clusters or distributed compute infrastructure at production scale - Hands-on work with GPU kernel optimization (CUTLASS, custom kernels, Nsight profiling) - Track record of building or running high-performance infrastructure for AI workloads (training or inference platforms) - Ability to reason from first principles and optimize for both memory-bound and compute-bound scenarios COMPENSATION AND BENEFITS: - $180,000 - $440,000 USD - Base salary is part of a total rewards package including equity, comprehensive medical, vision, and dental coverage, 401(k) retirement plan, short and long-term disability insurance, life insurance, and additional discounts and perks. SpaceXAI is an equal opportunity employer. For details on data processing, view the Recruitment Privacy Notice.
Technology

xAI
Backend Engineer - API
Senior
On-site
Palo Alto, CA
🏢 Summary: Engineering role focused on building and operating a highly scalable, low-latency API infrastructure that serves AI models globally. The position involves owning end-to-end distributed systems for high-throughput inference, including model serving, request routing, and observability. It requires deep expertise in Rust or C++ and strong experience with distributed systems and production-grade infrastructure. 🗂️ Requirements: Expert knowledge of Rust or C++, Experience designing and maintaining horizontally scalable distributed systems, Experience building reliable production infrastructure, Knowledge of service observability and reliability best practices, Experience operating PostgreSQL, Clickhouse, or MongoDB, Strong understanding of high-throughput, low-latency systems 📃 Skills: Rust, C++, Go, PostgreSQL, Clickhouse, MongoDB, gRPC, Docker, Kubernetes, TensorRT, vLLM, SGLang 🏢 Description: ABOUT THE ROLE: As an ideal candidate you have a good understanding of how highly scalable and reliable production infrastructure is built. Most of our backend infrastructure is written in Rust. Familiarity with a compiled language such as C++, Rust, or Go is highly beneficial. RESPONSIBILITIES: Build the xAI API that serves our models to developers worldwide Own the end-to-end system responsible for high-throughput inference, handling billions of tokens per minute with low latency and high availability, including model serving infrastructure, request routing, SDK development, rate limiting, observability, and efficient scaling BASIC QUALIFICATIONS: Expert knowledge of either Rust or C++ Experience in designing, implementing, and maintaining reliable and horizontally scalable distributed systems Knowledge of service observability and reliability best practices Experience in operating commonly used databases such as PostgreSQL, Clickhouse, and MongoDB PREFERRED SKILLS AND EXPERIENCE: Experience with LLM inference engines and serving frameworks (e.g., SGLang, TensorRT, vLLM) Experience designing or building with agent SDKs and agent orchestration frameworks Experience with Docker, Kubernetes, and containerized applications Expert knowledge of gRPC (unary, response streaming, bi-directional streaming, REST mapping) COMPENSATION AND BENEFITS $180,000 - $440,000 USD Base salary is just one part of the total rewards package, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short- and long-term disability insurance, life insurance, and various other discounts and perks. xAI is an equal opportunity employer. For details on data processing, view the Recruitment Privacy Notice.
Technology

xAI
Software Engineer - Training/Inference (C++)
Senior
On-site
Palo Alto, CA
🏢 Summary: High-impact inference engineering role focused on building and optimizing large-scale distributed model serving systems for Grok. The position involves low-level GPU and inference optimization, scalable infrastructure development, and ensuring high-performance, reliable AI serving at massive scale. Candidates will work across the full inference stack, from orchestration to GPU kernels and CI/CD systems. 🗂️ Requirements: Deep systems programming in C/C++ or Rust, Experience with large-scale high-concurrency production serving, Experience with GPU inference engines, Strong knowledge of batching, caching, load balancing, and parallelism, Experience with GPU kernel optimization and code generation, Knowledge of quantization, speculative decoding, distillation, and low-precision numerics, Experience with testing, benchmarking, and reliability engineering for inference services, Experience designing and implementing CI/CD infrastructure for inference 📃 Skills: C, C++, Rust, vLLM, SGLang, Triton, TensorRT-LLM, GPU, CI/CD, Quantization, Distillation, Parallelism, Caching, Benchmarking, Load-balancing, Autoscaling 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. About the Role: We are building the high-performance inference platform that serves Grok to millions of users every day with lightning speed and perfect reliability. As a Member of Technical Staff - Inference, you will design and optimize large-scale model serving systems end-to-end. You will own everything from distributed infrastructure (global KV cache, continuous batching, load balancing, auto-scaling) to deep low-level optimizations (GPU kernels, quantization, speculative decoding, tail latency). This is a high-impact role where your work directly determines how fast and reliably users interact with Grok at massive scale. Responsibilities: - Architect and implement scalable distributed infrastructure for model serving (load balancing, auto-scaling, batch scheduling, global KV cache). - Optimize latency and throughput of model inference under real production workloads. - Build reliable, high-concurrency serving systems that serve billions of users with 100% uptime, 0% error rate, and excellent tail latency. - Benchmark, fine-tune, and accelerate inference engines (including low-level GPU kernel work and code generation). - Develop custom tools to trace, replay, and fix issues across the full stack — from orchestration down to GPU kernels. - Create robust CI/CD infrastructure for seamless endpoint deployment, image publishing, and inference engine updates. - Accelerate research on scaling test-time compute, RL rollout, and model-hardware co-design for next-generation systems. Basic Qualifications: - Deep low-level systems programming (C/C++ or Rust) - Experience with large-scale, high-concurrent production serving. - Experience with GPU inference engines (vLLM, SGLang, Triton, TensorRT-LLM, etc.). - Strong background in system optimizations: batching, caching, load balancing, parallelism. - Low-level inference optimizations: GPU kernels, code generation. - Algorithmic inference optimizations: quantization, speculative decoding, distillation, low-precision numerics. - Experience with testing, benchmarking, and reliability of inference services. - Experience designing and implementing CI/CD infrastructure for inference. Compensation and Benefits: - $180,000 - $440,000 USD - Base salary is just one part of the total rewards package, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks. SpaceXAI is an equal opportunity employer. For details on data processing, view the Recruitment Privacy Notice.
Technology

xAI
Backend Engineer - API
Senior
On-site
Palo Alto, CA
🏢 Summary: Engineering role focused on building and owning a high-throughput, low-latency API and backend infrastructure for large-scale model inference. The position involves designing and operating reliable, horizontally scalable distributed systems that serve billions of tokens per minute. You will develop and maintain model serving, routing, SDKs, and observability within a production-grade environment. 🗂️ Requirements: Expert knowledge of Rust or C++, Experience building and maintaining horizontally scalable distributed systems, Experience designing reliable high-availability production infrastructure, Knowledge of observability and reliability best practices, Experience operating PostgreSQL, Clickhouse, or MongoDB 📃 Skills: Rust, C++, Go, PostgreSQL, Clickhouse, MongoDB, gRPC, Docker, Kubernetes, TensorRT, vLLM, SGLang, REST, SDK 🏢 Description: About xAI xAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.ABOUT THE ROLE: As an ideal candidate you have a good understanding of how highly scalable and reliable production infrastructure is built. Most of our backend infrastructure is written in Rust. So familiarity with a compiled language such as C++, Rust, or Go is highly beneficial. RESPONSIBILITIES: Build the xAI API that serves our models to developers worldwide Own the end-to-end system responsible for high-throughput inference, handling billions of tokens per minute with low latency and high availability, including model serving infrastructure, request routing, SDK development, rate limiting, observability, and efficient scaling BASIC QUALIFICATIONS: Expert knowledge of either Rust or C++ Experience in designing, implementing, and maintaining reliable and horizontally scalable distributed systems Knowledge of service observability and reliability best practices Experience in operating commonly used databases such as PostgreSQL, Clickhouse, and MongoDB PREFERRED SKILLS AND EXPERIENCE: Experience with LLM inference engines and serving frameworks (e.g., SGLang, TensorRT, vLLM) Experience designing or building with agent SDKs and agent orchestration frameworks Experience with Docker, Kubernetes, and containerized applications Expert knowledge of gRPC (unary, response streaming, bi-directional streaming, REST mapping) COMPENSATION AND BENEFITS $180,000 - $440,000 USD Base salary is just one part of our total rewards package at xAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.xAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.
Technology
Xebia sp. z o.o.
👉 Senior Platform Engineer
Senior
Remote
Wroclaw, Poland
20,000 - 31,000 PLN
🏢 Summary: The offer is for a Senior Platform Engineer responsible for designing, building, and operating scalable Kubernetes-based platforms on AWS. The role focuses on backend platform development, Infrastructure as Code with Terraform, CI/CD automation, and improving developer experience through self-service and multi-tenant solutions. It involves driving reliability, observability, and automation across a cloud-native ecosystem. 🗂️ Requirements: 5+ years of experience as a Platform Engineer, Strong hands-on experience with Kubernetes, including cluster operations, Strong programming skills in TypeScript or Go, Experience with AWS or other cloud platforms, Experience with Infrastructure as Code tools, preferably Terraform, Experience with CI/CD pipelines and deployment automation, Experience designing scalable multi-tenant systems, Strong understanding of API and platform architecture, Solid Linux and networking fundamentals, Experience with observability and monitoring solutions, Experience building self-service platform capabilities, Practical experience using AI-powered coding assistants, Work permit and residence in the European Union 📃 Skills: Kubernetes, AWS, Terraform, TypeScript, Go, Python, CI/CD, Linux, Networking, APIs, Docker, OpenTelemetry, Backstage, GitHub, Claude, Cursor 🏢 Description: 🟣 You will be: designing, building, and operating scalable Kubernetes-based platform infrastructure, developing backend platform services and reusable engineering components, building and maintaining cloud-native solutions (primarily on AWS), creating and evolving Infrastructure as Code solutions using Terraform, designing and implementing CI/CD pipelines and deployment automation, developing scalable multi-tenant platform capabilities for engineering teams, building self-service developer platform features and internal tooling, designing and maintaining APIs and platform architecture standards, improving platform reliability, observability, and monitoring capabilities, collaborating with engineering teams to improve developer experience and platform adoption, contributing to platform engineering best practices and technical standards, supporting container platform operations and orchestration processes, driving automation and operational excellence across the platform ecosystem. 🟣 Your profile: 5+ years of experience as a Platform Engineer, practical experience using AI-powered assistants (e.g. Claude Code, GitHub Copilot, Cursor) to improve productivity, quality, or decision-making in software delivery, strong hands-on experience with Kubernetes, including cluster operations and orchestration concepts, strong backend engineering background, strong programming skills in TypeScript and/or Go; Python experience is a plus, solid experience with cloud platforms, preferably AWS, practical experience with Infrastructure as Code tools, preferably Terraform, experience with CI/CD pipelines and deployment orchestration, experience designing scalable multi-tenant systems, strong understanding of API and platform architecture, solid Linux and networking fundamentals, experience with observability and monitoring solutions, experience building reusable self-service platform capabilities for developers, good English communication skills (at least B2 level). 🟣 Nice to have: experience applying GenAI in a more structured way within the SDLC, including defined workflows, prompt patterns, or tool integrations embedded into daily work, interest in and familiarity with emerging AI-driven practices (e.g. agent-based workflows, automation patterns, AI-augmented development), with a willingness to explore and experiment beyond standard approaches, experience with Internal Developer Platforms (IDP), experience building Kubernetes operators and controllers, experience with Backstage or developer portal tooling, experience in container runtime or platform engineering, knowledge of event-driven and distributed systems architecture, experience with service mesh technologies and OpenTelemetry. Work from the European Union region and a work permit are required. 🟣 Recruitment Process: CV review – HR Call – Interview – Client Interview – Decision 🎁 Benefits 🎁 ✍ Development: development budget of up to 6,800 PLN, we fund certifications e.g.: AWS, Azure, ISTQB, PSM, access to Udemy, Safari Books Online and more, events and technology conferences, technology Guilds, internal training, Xebia Library, Xebia Upskill. 🩺 We take care of your health: private medical healthcare, multiSport card - we subsidise a MultiSport card, mental Health Support. 🤸♂️ We are flexible: flexible working hours, B2B or permanent contract, contract for an indefinite period.
Technology

SpaceXAI
Software Engineer - Training/Inference (C++)
Senior
On-site
Palo Alto, CA
129,996 - 234,996 USD/yr
🏢 Summary: High-impact engineering role focused on building and optimizing large-scale AI inference systems for serving Grok at massive scale. The position involves distributed infrastructure, GPU-level optimization, inference acceleration, and reliability engineering for high-concurrency production environments. Candidates will work end-to-end on scalable model serving, CI/CD infrastructure, and next-generation inference optimization. 🗂️ Requirements: Deep low-level systems programming in C/C++ or Rust, Experience with large-scale high-concurrency production serving, Experience with GPU inference engines, Strong background in batching, caching, load balancing, and parallelism, Experience with GPU kernels and code generation, Knowledge of quantization, speculative decoding, distillation, and low-precision numerics, Experience with testing, benchmarking, and reliability of inference services, Experience designing and implementing CI/CD infrastructure for inference 📃 Skills: C++, Rust, vLLM, SGLang, Triton, TensorRT-LLM, GPU, CI/CD, Quantization, Distillation, Caching, Batching, Parallelism, Load-balancing, Autoscaling 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. About the Role: We are building the high-performance inference platform that serves Grok to millions of users every day with lightning speed and perfect reliability. As a Member of Technical Staff - Inference, you will design and optimize large-scale model serving systems end-to-end. You will own everything from distributed infrastructure (global KV cache, continuous batching, load balancing, auto-scaling) to deep low-level optimizations (GPU kernels, quantization, speculative decoding, tail latency). This is a high-impact role where your work directly determines how fast and reliably users interact with Grok at massive scale. Responsibilities: - Architect and implement scalable distributed infrastructure for model serving (load balancing, auto-scaling, batch scheduling, global KV cache). - Optimize latency and throughput of model inference under real production workloads. - Build reliable, high-concurrency serving systems that serve billions of users with 100% uptime, 0% error rate, and excellent tail latency. - Benchmark, fine-tune, and accelerate inference engines (including low-level GPU kernel work and code generation). - Develop custom tools to trace, replay, and fix issues across the full stack — from orchestration down to GPU kernels. - Create robust CI/CD infrastructure for seamless endpoint deployment, image publishing, and inference engine updates. - Accelerate research on scaling test-time compute, RL rollout, and model-hardware co-design for next-generation systems. Basic Qualifications: - Deep low-level systems programming (C/C++ or Rust) - Experience with large-scale, high-concurrent production serving - Experience with GPU inference engines (vLLM, SGLang, Triton, TensorRT-LLM, etc.) - Strong background in system optimizations: batching, caching, load balancing, parallelism - Low-level inference optimizations: GPU kernels, code generation - Algorithmic inference optimizations: quantization, speculative decoding, distillation, low-precision numerics - Experience with testing, benchmarking, and reliability of inference services - Experience designing and implementing CI/CD infrastructure for inference Compensation and Benefits: - $180,000 - $440,000 USD - Equity compensation - Medical, vision, and dental coverage - 401(k) retirement plan - Short and long-term disability insurance - Life insurance - Additional discounts and perks SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.
Technology

SpaceXAI
Member Of Technical Staff - Cloud Infrastructure
Senior
On-site
Palo Alto, CA
180,000 - 440,004 USD/yr
🏢 Summary: Senior Infrastructure Engineer role focused on building and operating secure, scalable AI infrastructure for US government projects across bare metal and classified cloud environments. The position involves Kubernetes-based infrastructure management, GPU cluster operations, observability, automation, and compliance-driven reliability engineering. This is an in-person role in Palo Alto or Washington, DC with up to 50% travel. 🗂️ Requirements: Active Top Secret (TS) security clearance, 5+ years of infrastructure or site reliability engineering experience, Experience building and maintaining scalable systems, Proficiency with Pulumi, Terraform, or Ansible, Deep knowledge of Kubernetes, CNI, CRI, and CSI, Experience with incident management and SLAs/SLOs, Strong communication and documentation skills, Ability to work in secure or government environments, Willingness to travel up to 50%, On-site availability in Palo Alto, CA or Washington, DC 📃 Skills: Kubernetes, CNI, CRI, CSI, Pulumi, Terraform, Ansible, GPU, Kyverno, ArgoCD, Go, IaC, Observability, SLA, SLO 🏢 Description: SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: We are seeking a highly skilled Senior Infrastructure Engineer to join our US Government Team, focused on designing, building, and operating secure, scalable infrastructure for critical government projects. In this role, you will develop and manage training and inference clusters, as well as highly reliable applications, across bare metal, classified cloud, and hybrid cloud architectures. You will leverage your expertise in Kubernetes and GPU hardware to deliver robust, secure systems that support large-scale AI workloads while meeting stringent federal compliance requirements. This role demands a passion for automation, observability, and ensuring system integrity in a fast-paced, high-security environment. RESPONSIBILITIES: - Develop and optimize software to provision and manage xAI's infrastructure across on-premise, virtual machine, and classified cloud environments, enabling efficient scaling for US government initiatives. - Enhance the reliability, performance, and cost-effectiveness of infrastructure to support large-scale AI and application workloads in secure, classified settings. - Collaborate with xAI engineers to understand workload requirements and design tailored solutions that meet government-specific needs and compliance standards. - Implement robust observability, monitoring, and security practices to ensure the integrity, availability, and confidentiality of critical systems, adhering to federal protocols. - Manage storage infrastructure using Infrastructure-as-Code (IaC) tools such as Pulumi, Terraform, or Ansible, with a focus on secure data handling. - Drive system reliability through incident management, postmortems, and the definition of clear SLAs and SLOs, while maintaining security and compliance. - This is an in-person role based in Palo Alto, CA or Washington, DC, with up to 50% travel required. BASIC QUALIFICATIONS: - Active Top Secret (TS) security clearance. - 5+ years of experience as an Infrastructure Engineer, Site Reliability Engineer, or similar role, with a focus on building and maintaining reliable, scalable systems, preferably in secure or government environments. - Proficiency in managing storage infrastructure with IaC tools such as Pulumi, Terraform, or Ansible. - Deep understanding of the Kubernetes stack, including CNI, CRI, CSI, and related components. - Demonstrated ability to improve system reliability through incident management, postmortems, and defining SLAs/SLOs. - Excellent communication and documentation skills, with the ability to handle sensitive information concisely and accurately. PREFERRED SKILLS AND EXPERIENCE: - Deep familiarity with installing and using GPU hardware, including setting up drivers, debugging issues, and ensuring reliability. - Experience with high-traffic web or mobile application workloads, including optimizing Kubernetes for large-scale deployments in classified or federal settings. - Familiarity with chaos engineering, capacity planning, or similar practices for ensuring system resilience in government projects. - Proficiency with tools such as Kyverno, ArgoCD, or Go programming for infrastructure automation. - Strong sense of ownership, curiosity, and enthusiasm for tackling complex technical challenges in secure environments. - Passion for problem-solving and a proactive drive to deliver impactful results while adhering to security protocols. - Certifications in security-related fields (e.g., CISSP) or experience in secure federal environments. COMPENSATION AND BENEFITS: - $180,000 - $440,000 USD - Base salary is just one part of the total rewards package, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks. SpaceXAI is an equal opportunity employer. For details on data processing, view the Recruitment Privacy Notice.
Technology
ITDS
Kubernetes & Cloud Infrastructure Engineer – AI Platform
Senior
On-site
Wroclaw, Poland
40,000 - 60,000 PLN
🏢 Summary: On-site role in Wroclaw for a Kubernetes & Cloud Infrastructure Engineer focused on building and operating scalable, secure AI platform infrastructure. The position centers on managing Kubernetes clusters, AWS environments, and infrastructure as code to support GPU workloads and large language model deployments. You will enhance reliability, automation, and observability of cloud-native AI systems. 🗂️ Requirements: 4+ years in Cloud, DevOps, or Platform Engineering, Strong expertise in Kubernetes cluster operations and multi-tenancy, Experience with GPU scheduling and model serving infrastructure, Hands-on experience with AWS (networking, IAM, EKS, EC2), Proficiency in Python, Experience with Terraform and infrastructure as code, Ability to build and operate CI/CD pipelines, Experience implementing GitOps workflows, Experience defining and managing infrastructure SLOs, Fluent English, Legal right to work in the EU 📃 Skills: Kubernetes, AWS, EKS, EC2, IAM, Terraform, Helm, Kustomize, Python, CI/CD, GitOps, GPU, Docker, Linux, Networking, Observability 🏢 Description: Unleash the power of cloud-native infrastructure — revolutionize AI platforms with your expertise! Wroclaw-based opportunity with on-site work model As a Kubernetes & Cloud Infrastructure Engineer – AI Platform , you will be working for our client, a leader in AI innovation, building and managing critical infrastructure that supports advanced AI models and large language model workloads. Your work will directly impact the delivery, scalability, and security of cutting-edge AI solutions, empowering teams across the firm to harness AI's full potential. Join us and be part of shaping the future of intelligent technologies! Your main responsibilities: Build and operate scalable Kubernetes clusters supporting multi-tenancy, GPU workloads, and model serving. Manage AWS infrastructure, including networking, IAM, security, and cost optimization. Develop and maintain infrastructure as code utilizing tools like Terraform, Helm, and Kustomize. Implement and maintain CI/CD and GitOps workflows to streamline deployment pipelines. Build observability solutions for system health, utilization, latency, and platform performance monitoring. Automate scaling, capacity management, and enforce security and governance policies across cloud and on-premise environments. Define and monitor infrastructure Service Level Objectives (SLOs) to ensure reliability and performance. Collaborate closely with AI Platform Engineers, Data Scientists, and Research teams to support model inference and deployment infrastructure. Enable internal teams through platform tooling, onboarding, and self-service portals. You're ideal for this role if you have: 4+ years of experience in Cloud, DevOps, or Platform Engineering. Strong expertise in Kubernetes, including cluster operations, multi-tenancy, GPU scheduling, and Helm/Kustomize. Hands-on experience with AWS cloud services—networking, IAM, EKS, EC2, and cost management. Proficient in Python and infrastructure-as-code tools such as Terraform. Proven ability to build and operate CI/CD and GitOps workflows. Skilled in defining and managing infrastructure SLOs. Pragmatic, ownership-driven approach and comfortable working in ambiguous environments. It is a strong plus if you have: Experience supporting LLM inference workloads, model serving platforms, or GPU-backed infrastructure. Knowledge of RAG systems, vector databases, or AI-related data platforms. Background in high-performance or latency-sensitive environments. Relevant certifications such as CKA, CKAD, Solutions Architect, or similar. Language Required for the role: Fluent English Eligibility to work in this role: Only candidates with an existing legal right to work in the European Union will be considered for this role. #MAKEYourCareerBETTER Interested? Apply now and include your CV (preferably in English) along with a statement confirming your consent to the processing and storage of your personal data.