Unlock Full Resume Report

New offer - be the first one to apply!

September 27, 2026

Cloud Hardware Development Engineer (AWS AI/ML UltraServers)

Mid

120,000 - 210,000 USD/yr

Seattle, WA

Quick Facts

  • Defines server hardware architecture for GPU-accelerated AI/ML training and inference
  • Owns hardware validation from PCBA bring-up through rack integration
  • Leads ODM/JDM partners and drives fleet quality post-launch

Description

Own the end-to-end hardware architecture and validation of GPU-accelerated AI/ML server platforms. You will translate workload and customer requirements into detailed component specifications, execute validation from first silicon through fleet-scale deployment, and triage failures across PCIe, power delivery, memory, and accelerator interconnects. You will monitor fleet quality metrics and feed root-cause findings back into design and process improvements.

Role Description / Responsibilities

  • Define server architectures based on workload demand and translate them into detailed designs and component specifications for high-performance AI workloads
  • Partner with interdisciplinary teams (component, firmware, test, qualification, integration) to deliver cohesive designs
  • Lead design reviews with ODM/JDM partners covering schematic, layout, BOM, and manufacturing DFx (design for test/manufacturing)
  • Define and execute validation strategies from PCBA bring-up through server and rack integration (power sequencing, signal integrity, thermal characterization, accelerator interconnect performance)
  • Own hardware debug during EVT/DVT/PVT builds and correlate failures across PCIe, power rails, memory channels, and GPU subsystems
  • Triage hardware issues at ODM facilities and datacenters, perform root cause analysis, and implement corrective actions
  • Own post-launch fleet quality metrics (annualized failure rates, component-level failure modes) and monitor operational telemetry for systemic issues
  • Work with test and automation teams to improve manufacturing yield and reduce test dwell times
  • Collaborate with EC2 architecture teams on instance definitions, workload requirements, and platform trade-offs; coordinate with firmware/software/operations to ensure designs are debuggable, serviceable, and automation-ready

Requirements

  • Bachelor’s degree in electrical engineering, computer engineering, or equivalent
  • Experience in server technologies (thermal, mechanical, power, signal integrity)
  • Experience developing functional specifications, design verification plans, and functional test procedures
  • 2+ years of hardware design, development, and validation experience for server or compute platforms

Benefits

  • Comprehensive health insurance (medical, dental, vision, prescription, Basic Life & AD&D, supplemental life options)
  • EAP and mental health support; medical advice line
  • Flexible Spending Accounts; adoption and surrogacy reimbursement coverage
  • 401(k) matching; paid time off; parental leave

Preferred Qualifications

  • Master’s degree in electrical engineering, computer engineering, or related field
  • Experience with analog, digital, and high-speed circuit design
  • Experience working in data centers or critical infrastructure
  • Experience with ODMs across product development and manufacturing lifecycle (2+ years)
  • Experience with hardware bring-up/debug/validation of GPU/accelerator platforms (2+ years)
  • Familiarity with PCIe topology, NVMe, and accelerator interconnects
  • Experience developing and executing test procedures for electrical or mechanical systems

Similar jobs you might like