Unlock Full Resume Report

New offer - be the first one to apply!

September 24, 2026

Senior Storage Software Engineer (DGX Cloud)

Senior

205,000 - 290,000 USD/yr

Santa Clara, CA

Quick Facts

  • Role: Senior Storage Software Engineer (hands-on technical lead) for DGX Cloud storage

  • Focus: Open-source parallel/distributed file systems and distributed object storage, production incident triage, performance/durability validation, and fleet configuration standards

Description

You will contribute to open-source parallel and distributed file systems and distributed object storage, and upstream fixes and features with community maintainers. As a hands-on storage software lead, you will write and review production code, read kernel/NFS/NVMe-oF/SPDK source when needed, and make technical calls against measurable storage delivery targets. You will triage and root-cause storage issues at tens of thousands of GPU scale, validate architecture with scale tests/benchmarks and recovery drills, and define tuning and operational best practices for high-performance GPU infrastructure.

Responsibilities

  • Contribute to open-source parallel/distributed file systems and distributed object storage; upstream fixes/features

  • Write and review production code; debug using kernel and storage stack sources

  • Triage and root-cause large storage incidents across GPU clusters (I/O/metadata, corruption, recovery)

  • Validate storage architecture and durability/performance via scale tests, benchmarks, and recovery drills

  • Recommend configuration, tuning, and operational guidelines for GPU-based file systems

  • Partner with training/inference, SRE/operations, networking, security, and storage/cloud/neocloud stakeholders

  • Use modern AI coding/agentic tools to accelerate building, debugging, validation, and operations

Requirements

  • BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field OR equivalent experience

  • 12+ years of direct storage software engineering experience, including multi-petabyte parallel/distributed file system work

  • Active open-source contributions to distributed/parallel file systems; write/review production code and personally conduct scale tests/recovery drills

  • Experience diagnosing/resolving storage problems in large GPU or HPC clusters with I/O and metadata performance analysis

  • Strong systems-language proficiency (C/C++/Rust/Go) and Python proficiency

  • Comfortable with Linux kernel storage/networking stacks (block layer, RDMA/RoCE/InfiniBand, NVMe, page cache, VFS, multipath)

  • Understanding of object storage (S3/Swift-class) and block storage (NVMe-oF, iSCSI)

  • Strong communication skills to clarify complex technical trade-offs

  • Comfortable with 24/7 production operations where storage incidents impact GPU availability; security-first approach

  • 100% hands-on engineering mindset (no delegation of debugging/tests)

Benefits

  • Base salary range for senior levels; eligibility for equity and benefits

  • Opportunity to build foundational storage capabilities for the AI era on large GPU fleets

Similar jobs you might like