Unlock Full Resume Report

New offer - be the first one to apply!

September 14, 2026

Director of Engineering, Flex Compute

Senior • On-site

285,000 - 335,000 USD/yr

Sunnyvale, CA

Quick Facts

Power is the binding constraint on AI infrastructure. The Flexible Compute system reduces data-center power draw on demand quickly, reliably, and verifiably while protecting workloads and riding through grid-stress events without breaking SLAs.

Description

You will operate at the seam between a fast-growing energy footprint and the software that controls it. This is a utility-facing, safety-critical control system built from the ground up, where you’ll define the architecture, ship the first production system, and build the team from scratch.

Responsibilities

  • Own the curtailment orchestration and decision layer: grid-signal ingestion, staged shedding, and per-SKU power capping without ever silently dropping a paid workload.
  • Land a utility-validated pilot: fast ramp to setpoint, tight accuracy, high-fidelity telemetry, and the test harness that proves it before utility commitment.
  • Deliver dynamic power management for oversubscription (more GPUs per megawatt) with no customer-visible impact.
  • Build fast, workload-aware GPU power estimation validated against fleet telemetry—used for shed forecasting, oversubscription admission, and power planning for new silicon.
  • Design for safety: authenticated signal ingress, bounded blast radius, fail-safe defaults, manual backstops.
  • Own graceful ride-through of power-loss and grid-stress events integrated with on-site battery and generation backstops.
  • Partner with Data Center Engineering and Energy teams on interconnection commitments, curtailment program design, and BESS/generation integration.
  • Integrate with the cloud control plane across Kubernetes and Slurm fleets (one control plane, not two).
  • Make build-vs-leverage calls across vendor power-management stacks and grid integration layers.
  • Hire and lead the team from the ground up, keeping headcount sublinear to fleet growth through automation.

Requirements

  • 12+ years in software engineering, including 5+ leading engineering teams; ideally took a system from 0→1 to production scale.
  • Deep experience with distributed control planes, orchestration, or fleet automation (e.g., Temporal, Kubernetes, Slurm).
  • Track record with safety-critical or physically-actuating systems where a misfire has a bounded, designed-for worst case.
  • Working knowledge of data-center power systems: utility interconnection, switchgear/UPS/BESS, rack/PDU distribution, power telemetry, and GPU power management.
  • Experience with energy markets or grid programs demand response and curtailable-load tariffs; ISO/RTO market signals (e.g., ERCOT, PJM, CAISO) or demonstrated speed to fluency.
  • Strong build-vs-buy judgment and comfort deciding with incomplete data.
  • Record of hiring senior engineers and running healthy operations for systems that must never fail silently.

Benefits

  • Competitive compensation and equity packages, including Restricted Stock Units.
  • Paid time off, paid holidays, and leave of absence programs.
  • Comprehensive health, dental, and vision insurance; employer contributions to an HSA account.
  • Paid parental leave; paid life insurance; short-term and long-term disability.
  • Professional development and tuition reimbursement; mental health and wellness support.
  • Commuter benefits (parking and transit) and a cell phone stipend.
  • 401(k) Retirement plan with company match up to 4% of salary.
  • Volunteer time off; global travel insurance & emergency assistance; daily meals allowance; additional perks specific to location.

Compensation Range

Compensation is paid in the range of up to $285,000 - $335,000 + Bonus, with Restricted Stock Units included in all offers. Final compensation is determined by knowledge, education, abilities, internal equity, and alignment with market data.

Similar jobs you might like