New offer - be the first one to apply!

October 9, 2026

Senior Hardware Reliability Engineer (Infrastructure Reliability & Quality)

Senior

146,000 - 210,000 USD/yr

Seattle, WA

Quick Facts

Senior Hardware Reliability Engineer focused on Infrastructure Reliability & Quality.

Description

Drive reliability risk identification, assessment, and mitigation for datacenter infrastructure equipment (e.g., air handling units, liquid cooling, chillers). Perform root cause analysis of critical equipment failures and lead continuous improvements to increase datacenter availability. Apply Physics-of-Failure, accelerated life testing, stress analysis (including thermal, electrical, chemical, and mechanical stresses), statistical techniques, and reliability modeling to evaluate risks across design, manufacture, deployment, and field sustaining.

Role Description

  • Proactively drive reliability risk identification, assessment, and mitigation for datacenter infrastructure equipment.
  • Perform root cause analysis of critical failures and drive corrective and preventive actions.
  • Collaborate with suppliers and internal partners on product specification, risk identification plans, qualification, and remediation.
  • Develop reliability metrics and models to optimize datacenter configurations for availability and cost.
  • Monitor field performance during sustaining, and improve maintenance and replacement recommendations based on reliability data.

Key Job Responsibilities

  • Drive Design for Reliability (DFR) methodology for new product designs.
  • Drive reliability/quality qualification of third-party critical infrastructure equipment for AWS data centers.
  • Oversee factory and site testing across Liquid Cooling, generator, chiller, air handler, and related LLE categories.
  • Guide and validate root cause analysis conclusions (internal teams, OEMs, and external laboratories) and ensure high testing/remediation standards.
  • Recommend infrastructure maintenance and equipment replacement based on reliability data.
  • Provide feedback to sourcing/procurement teams on vendor performance.
  • Analyze internal reliability data and create metrics to maximize reliability at lowest cost.
  • Support DFMEAs as needed.
  • Develop end-of-life strategy for critical infrastructure equipment.

Requirements

  • Experience in industrial or commercial engineering in mission-critical facilities, including data centers, power generation, or oil and gas facilities.
  • Bachelor’s or Master’s degree in Reliability Engineering, Physics, Mechanical or Materials Engineering, or related field.
  • 6+ years of Reliability Engineering experience in a high-reliability industry.
  • 3+ years experience with accelerated life testing, stress analysis, and finite element analysis.
  • Ability to travel within the US and internationally.

Benefits

  • Sign-on payments and restricted stock units (RSUs) as part of compensation.
  • Comprehensive health benefits (medical, dental, vision, prescription, Basic Life & AD&D; optional Supplemental Life), EAP, and mental health support.
  • Medical Advice Line, Flexible Spending Accounts, and adoption/surrogacy reimbursement coverage.
  • 401(k) matching, paid time off, and parental leave.

Similar jobs you might like

Unlock Full Resume Report