Unlock Full Resume Report

This position is no longer accepting applications

Positions open for more than 30 days are automatically closed and marked as expired

Don't let one closed door slow you down, here's your next move:

May 22, 2026

Sr. Site Reliability Engineer (SRE)

Senior • Remote

165,000 - 225,000 USD/yr

Moonlite delivers high-performance AI infrastructure for organizations running intensive computational research, large-scale model training, and demanding data processing workloads.We provide infrastructure deployed in our facilities or co-located in yours, delivering flexible on-demand or reserved compute that feels like an extension of your existing data center. Our team of AI infrastructure specialists combines bare-metal performance with cloud-native operational simplicity, enabling research teams and enterprises to deploy demanding AI workloads with enterprise-grade reliability and compliance.

Your Role:

You will be instrumental in building and operating production-grade AI infrastructure with deep Kubernetes expertise at its core. Working closely with our systems engineers, network engineers, and platform engineering team, you'll architect and operate the Kubernetes infrastructure that powers our control plane and orchestrates compute, storage, and networking at scale. This role requires deep understanding of Kubernetes internals, custom resource definitions (CRDs), storage and network integrations, and building production-grade clusters from the ground up (not just deploying in managed environments). You'll ensure enterprise-grade reliability while establishing the automation, observability, and operational practices.

Job Responsibilities

  • Kubernetes Infrastructure Engineering: Design, build, and operate production Kubernetes clusters on bare-metal infrastructure – including cluster bootstrapping, control plane architecture, etcd management, and scaling strategies for high-performance compute workloads.
  • Kubernetes Networking & CNIs: Implement and operate custom Kubernetes networking solutions with SR-IOV for high-performance GPU interconnects, multi-tenancy isolation and advanced networking policies. Configure CNI plugins and network segmentation for research workloads.
  • Custom Operators & Controllers: Develop and maintain custom Kubernetes operators and controllers for bare-metal provisioning, infrastructure lifecycle management, and resource orchestration across compute, storage, and networking domains.
  • GPU Infrastructure Integration: Deploy and optimize NVIDIA GPU operators, device plugins, and other custom scheduling logic for GPU workload placement and utilization optimization.
  • Platform Integration & Storage: Build deep integrations between Kubernetes and underlying infrastructure including CSI drivers for storage, custom admission controllers for policy enforcement, and scheduling extensions for specialized hardware placement.
  • Infrastructure Automation: Design and implement automation using Terraform, Ansible, Helm, and custom operators to orchestrate infrastructure workflows and enable deployments across multiple regions.
  • Production Operations & Reliability: Manage production bare-metal infrastructure across multiple regions. Build systems ensuring high availability, fault tolerance, and graceful degradation – establishing SLIs, SLOs, and monitoring to meet enterprise reliability commitments.
  • Observability & Incident Response: Build comprehensive monitoring, logging, and alerting using Prometheus, Grafana, and ELK stack. Lead incident response, conduct postmortems, and implement preventative measures to improve reliability and reduce MTTR.
  • Performance & Capacity Planning: Identify and resolve performance bottlenecks across infrastructure domains. Monitor utilization trends, forecast capacity needs, and optimize resource allocation for various workloads.

Requirements

  • Experience: 5+ years in SRE, DevOps, or infrastructure engineering roles with proven experience operating production infrastructure at scale.
  • Kubernetes Infrastructure Expertise: Deep hands-on experience building and operating production Kubernetes clusters on bare-metal infrastructure – not just deploying workloads in managed clusters. Must understand cluster bootstrapping, control plane architecture, etcd operations, and scaling strategies.
  • Kubernetes Internals & Integration: Strong understanding of Kubernetes internals including custom resource definitions (CRDs), operators, controllers, admission webhooks, and scheduling. Experience integrating storage (CSI drivers), networking (CNI, SR-IOV), and specialized hardware (GPU device plugins) with Kubernetes.
  • Linux Systems Experience: Strong fundamentals in Linux systems administration, performance tuning, troubleshooting, and automation in production environments.
  • Infrastructure Automation: Proficiency with infrastructure-as-code tools (Terraform, Ansible, Helm) and building automation to reduce operational overhead.
  • Networking Fundamentals: Solid understanding of networking concepts including IPAM, DNS, DHCP, VLAN/VXLAN, routing, load balancing, and experience troubleshooting network issues in production.
  • Observability & Monitoring: Experience building and maintaining comprehensive monitoring solutions using tools like Prometheus, Grafana, and centralized logging systems.
  • Reliability Practices: Understanding of SRE principles including SLIs/SLOs/SLAs, error budgets, incident management, and blameless postmortems.
  • Scripting & Automation: Strong scripting skills in Go, Python, or Bash for automation, tooling development, and operational efficiency.
  • Problem-Solving Under Pressure: Demonstrated ability to troubleshoot complex issues under pressure, manage incidents effectively, and communicate clearly during outages.
  • Collaboration & Communication: Excellent communication skills and ability to work across teams including systems engineers, network engineers, and software developers.

Preferred Qualifications

  • Experience building custom Kubernetes operators or controllers for infrastructure orchestration
  • Deep familiarity with Kubernetes networking (Calico, Cilium, Multus), service mesh technologies, and network policy management
  • Experience with GPU workload orchestration including NVIDIA GPU Operator, MIG, time-slicing, and device plugins
  • Background with advanced Kubernetes features including custom schedulers, admission controllers, and API server extensions
  • Experience with Kubernetes cluster federation or multi-cluster management
  • Knowledge of high-performance networking technologies (InfiniBand, RDMA, RoCE) and their integration with Kubernetes
  • Experience with enterprise storage systems (VAST, Lightbits, Ceph, or similar)
  • Familiarity with configuration management at scale and GitOps practices
  • Understanding of security best practices for Kubernetes and bare-metal infrastructure
  • Experience operating infrastructure in regulated industries or co-located data center environments
  • Background supporting research institutions, technical computing environments, or enterprise AI infrastructure

Key Technologies

  • Kubernetes, Linux, Terraform, Ansible, Prometheus, Grafana, ELK Stack, Go, Python, Bash, NVIDIA GPU Technologies, High-Performance Networking, Enterprise Storage Systems

Why Moonlite

  • Build Critical Research Infrastructure: Your work will directly enable quantitative research teams and AI practitioners to push the boundaries of what's possible in financial modeling and AI research.
  • Enterprise Impact: Build and operate infrastructure that supports mission-critical research and AI workloads for leading financial institutions and research organizations.
  • Technical Excellence: Join an infrastructure team focused on delivering enterprise-grade reliability while pushing the boundaries of high-performance computing capabilities.
  • Hands-On Ownership: As part of our growing infrastructure team, you'll have significant ownership over critical systems and the autonomy to influence our operational practices and technology choices.
  • Industry Leadership: Work alongside experienced infrastructure professionals who have built and operated systems for the most demanding computing environments.

We offer a competitive total compensation package combining a competitive base salary, startup equity, and industry-leading benefits. The total compensation range for this role is $165,000 – $225,000, which includes both base salary and equity. Actual compensation will be determined based on experience, skills, and market alignment. We provide generous benefits, including a 6% 401(k) match, fully covered health insurance premiums, and other comprehensive offerings to support your well-being and success as we grow together.

#li-remote

Similar jobs you might like

Technology

New offer

Nebius

Forward Deployment Engineering Manager

Senior

Remote

225,800 - 281,000 USD/yr

🏢 Summary: Lead and grow a Forward Deployed Engineering team delivering production-quality AI integrations, reference architectures, and proofs of concept for an AI cloud platform. The role combines people leadership, technical architecture oversight, partner engagement, and hands-on prototyping across agentic AI, inference, infrastructure, and data systems. 🗂️ Requirements: 8+ years of hands-on AI application, ML systems, or AI infrastructure engineering experience, 2+ years of engineering team leadership or management experience, Deep practical knowledge of LLM APIs, inference runtimes, orchestration frameworks, vector databases, RAG architectures, and agentic pipelines, Hands-on experience with agentic frameworks, Strong Python programming and end-to-end AI prototyping skills, Experience defining reference architectures and technical patterns, Experience building API and developer-platform integrations, Ability to work with external partner engineering teams and internal product and engineering teams, Strong technical communication skills, Authorization to work in the country of application 📃 Skills: Python, LangChain, LangGraph, CrewAI, AutoGen, LLM, RAG, FastAPI, Flask, Docker, Kubernetes, Git, AWS, GCP, Azure, vLLM, SGLang, TensorRT-LLM, Transformers, Qdrant, Weaviate, Milvus, pgvector, CUDA, TensorRT, NeMo 🏢 Description: About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. About Nebius Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role Nebius builds the infrastructure serious AI teams run on - GPU clusters, inference runtimes, agent development environments, data pipelines - all of it purpose-built for the most demanding AI workloads. What we are now building is the ecosystem function that ensures the best AI companies choose to build on us, integrate with us, and stay. As Manager of Forward Deployed Engineering, you will lead a team of FDEs who sit at the intersection of solution architecture and hands-on engineering. You will set the technical bar, develop your team's craft, and ensure the work your team ships - integrations, reference architectures, and proofs of concept - is production-quality and strategically sound. You will also stay technically engaged: reviewing architectures, unblocking hard scoping problems, and occasionally prototyping yourself when the situation calls for it. You're welcome to work remotely in the United States. Your Responsibilities Will Include People & Team Leadership Hire, develop, and retain a team of Forward Deployed Engineers across ecosystem focus areas: agentic, inference, infrastructure, and data. Set clear expectations for technical quality and partner engagement; coach engineers toward those standards. Run structured 1:1s, provide direct and actionable feedback, and own the growth of each person on your team. Build a team culture where moving fast and building well are not in tension. Partner with Recruiting to define what great looks like for FDE roles and actively source candidates from the AI builder community. Technical Oversight & Architecture Review and elevate the technical work your team produces - integration architectures, proofs of concept, reference patterns, and partner scoping assessments. Serve as a technical escalation point for complex partner engagements; step in hands-on when needed. Maintain a high bar for what goes into the reference architecture library - not just what works, but what should be emulated. Stay current with the AI tooling ecosystem so you can guide your team's technical judgment, not just their output. Ecosystem Presence Represent Nebius at hackathons, in open source communities, and at technical events. Build in public - demos, reference architectures, and integrations that establish Nebius as the platform serious AI builders choose. Stay current with the AI tooling ecosystem - you know what shipped last week and what it means for our stack. Platform focus areas Depending on your background and mutual fit, you will focus on one or more of the following: Agentic - agent frameworks, memory systems, tool integration, orchestration, MCP, and guardrails Managed Inference - inference runtimes, model serving, optimization tooling, speculative decoding, and KV-cache routing IaaS / Managed Infrastructure - cloud-native integrations, GPU orchestration, and enterprise platform connectors Data - vector databases, retrieval systems, RAG architectures, data pipeline integrations, and synthetic data tooling Partner & Internal Stakeholder Engagement Engage directly with senior partner engineering leaders and founding CTOs; your team handles the working level, and you handle the strategic level. Translate what your team is seeing in the field into actionable product requirements for Nebius platform teams. Represent the FDE function in platform planning discussions as the technical voice of ecosystem integration. Work with ISV, SI, and field teams to scale solution adoption and drive revenue once integrations are ready. We expect you to have 8+ years of hands-on engineering experience in AI application development, ML systems, or AI infrastructure. 2+ years managing or leading a team of engineers, with a track record of developing technical talent. Deep working knowledge of the AI developer stack - LLM APIs, inference runtimes, orchestration frameworks, vector databases, RAG architectures, and agentic pipelines - built through shipping, not reading. Hands-on experience with agentic frameworks such as LangChain, LangGraph, CrewAI, AutoGen, or equivalent. Strong Python programming skills and comfort prototyping end-to-end AI systems quickly. Experience defining reference architectures and technical patterns - not just implementing them. Proven ability to move from idea to working prototype fast - you have shipped meaningful things under time pressure and found it energizing. Experience building integrations across APIs and developer platforms - you understand where the complexity actually lives. Comfort working across external partner engineering teams and internal Nebius Product and Engineering teams simultaneously. Strong technical communication - you can explain architecture decisions and integration findings to a founding CTO and a non-technical partner lead in the same day. It will be an added bonus if you have Experience with inference frameworks and optimization: vLLM, SGLang, TensorRT-LLM, speculative decoding, quantization, batching, and KV-cache routing. Familiarity with NVIDIA's software stack: CUDA, TensorRT, NeMo, or equivalent. Experience with multimodal AI models - vision-language, speech, or structured data. Won or placed at major AI hackathons in the past 12 months. Worked as a developer advocate, solutions engineer, or technical partner manager at a leading AI platform or developer tooling company. Been an early engineer at a YC-backed AI startup - you built the product under real constraints. Open source projects or public demos with meaningful community adoption. Proficiency with DevOps tools: Docker, Kubernetes, and Git. Preferred Technical Stack Languages - Python ML frameworks - vLLM, SGLang, TensorRT-LLM, Transformers, and OpenAI / Anthropic SDKs Agentic frameworks - LangChain, LangGraph, CrewAI, AutoGen, smolagents, or equivalent Vector databases - Qdrant, Weaviate, Milvus, and pgvector API and web frameworks - FastAPI and Flask DevOps - Kubernetes, Docker, and Git Cloud platforms - AWS, GCP, and Azure Key Employee Benefits Health Insurance: 100% company-paid medical, dental, and vision coverage for employees and families. 401(k) Plan: Up to 4% company match with immediate vesting. Parental Leave: 20 weeks paid for primary caregivers, 12 weeks for secondary caregivers. Remote Work Reimbursement: Up to $85/month for mobile and internet. Disability & Life Insurance: Company-paid short-term, long-term, and life insurance coverage. Pay Transparency We offer competitive compensation and benefits packages. Actual compensation will be determined based on job-related factors, including experience, skills, qualifications, the level at which the candidate is hired, and geographic location, consistent with applicable law. Starting Base Compensation Range: $225,800 - $281,000 USD Benefits & Perks Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know Pay Transparency We offer competitive compensation and benefits packages. Actual compensation will be determined based on job-related factors, including experience, skills, qualifications, the level at which the candidate is hired, and geographic location, consistent with applicable law. Base Compensation Range $225—$280 USD Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's It Like To Work At Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.