October 3, 2026

Cloud Hardware Development Engineer (Cloud AI/ML Server Platforms)

Mid

120,000 - 210,000 USD/yr

Seattle, WA

Quick Facts

  • Cloud Hardware Development Engineer for Cloud AI/ML server platforms
  • Owns end-to-end NPI to production fleet reliability for accelerator (AI/ML/GPU) servers running at scale

Description

Own the end-to-end development lifecycle for accelerator (AI/ML/GPU) server platforms—from NPI architecture and qualification through fleet health in production within the AWS cloud environment. You will partner with internal and external teams to design, validate, manufacture, launch, and continuously improve reliability using telemetry-driven diagnostics and automation. When complex failures occur during qualification or in large-scale fleets, you drive root cause analysis across hardware, firmware, software, and physical layers.

Responsibilities

NPI (New Product Introduction)

  • Own end-to-end NPI for storage and/or accelerator server platforms, from architecture definition through design verification, manufacturing ramp, and launch
  • Lead technical solutions for complex server and rack architectural challenges
  • Work with ODM/manufacturing partners to develop, validate, and manufacture server products at scale
  • Create functional specifications, design verification plans, and test procedures
  • Drive qualification and readiness milestones to meet performance, reliability, and cost targets
  • Identify and resolve technical risks early to prevent issues reaching production

Fleet Health, Diagnostics & Automation

  • Own fleet health for launched server platforms; reliability continues after shipment
  • Design predictive failure detection systems using telemetry, sensor data, error trending, and log correlation
  • Drive toward zero-touch operations for detection, diagnosis, and remediation without human intervention
  • Debug complex, time-sensitive system failures
  • Perform root cause analysis correlating firmware, kernel, driver, thermal, power, and physical layers

Systems Design & Technical Depth

  • Apply expertise across compute, storage, network, GPU/accelerator systems, processes, and operations
  • Design and implement solutions for system-level issues at large scale
  • Decompose complex server problems (testability, reliability, diagnostics) into deliverable tasks and features

Cross-Team Collaboration

  • Collaborate with hardware, software, manufacturing, supply chain, and product management teams
  • Ensure server hardware meets data path and control path requirements
  • Coordinate with datacenter operations to close the loop between field failures and design improvements

Benefits

  • Comprehensive health insurance (medical, dental, vision, prescription) and Basic Life & AD&D; option for Supplemental life plans
  • EAP and mental health support, Medical Advice Line
  • Flexible Spending Accounts
  • 401(k) matching
  • Paid time off
  • Parental leave
  • Sign-on payments and restricted stock units (RSUs)
  • Opportunity to improve performance, quality, and cost in an intellectually challenging, fast-paced environment

Similar jobs you might like

Unlock Full Resume Report