New offer - be the first one to apply!

October 9, 2026

Senior Hardware Reliability Engineer (Infrastructure Reliability & Quality)

Senior

160,000 - 220,000 USD/yr

Herndon, VA

Quick Facts

Senior Hardware Reliability Engineer focused on Infrastructure Reliability & Quality for datacenter equipment.

Description

Drive reliability risk identification, assessment, and mitigation for datacenter infrastructure equipment from design and manufacturing through deployment and sustaining. Lead root cause analysis of critical failures and drive continuous improvements to increase datacenter availability. Partner with suppliers and internal teams to qualify equipment, validate remediation/testing, and establish reliability/quality metrics using physics-of-failure and statistical approaches.

Responsibilities

  • Drive DFR (Design for Reliability) to design reliability into new product designs.
  • Perform reliability/quality qualification of third-party critical infrastructure equipment for datacenter use.
  • Oversee factory and site testing across categories (e.g., liquid cooling, generator, chiller, air handler).
  • Guide and validate root cause analyses for field failures (internal teams, OEMs, and external labs) and ensure high testing/remediation standards.
  • Recommend infrastructure maintenance and equipment replacement based on reliability data.
  • Provide feedback to sourcing/procurement on vendor performance.
  • Analyze internal reliability data and create metrics to improve reliability at lowest cost.
  • Support DFMEAs as needed.
  • Develop end-of-life strategy for critical infrastructure equipment.

Requirements

  • Experience in industrial or commercial engineering in mission-critical facilities (data centers, power generation, or oil and gas facilities).
  • Bachelor’s or Master’s degree in Reliability Engineering, Physics, Mechanical or Materials Engineering, or related field.
  • 6+ years of Reliability Engineering experience in a high-reliability industry.
  • 3+ years experience with accelerated life testing, stress analysis, and finite element analysis.
  • Experience with Physics-of-Failure-based analytical and empirical approaches across design, manufacture, and deployment.
  • Ability to conduct lifecycle environmental and operational stress-driven risk analysis (thermal, electrical, chemical, mechanical).
  • Ability to assess electronics manufacturing process-related quality/reliability issues.
  • Knowledge of statistical techniques/models to analyze test and field data.
  • Ability to develop datacenter system-level reliability models and perform reliability quantification and risk analysis.
  • Ability to perform sustaining activities: field performance monitoring, RCA, CAPA, and vendor auditing/reviews.
  • Ability to travel within the US and internationally.

Benefits

  • Comprehensive health insurance (medical, dental, vision, prescription, Basic Life & AD&D, and optional supplemental life), EAP, mental health support, medical advice line, flexible spending accounts, adoption and surrogacy reimbursement.
  • 401(k) matching.
  • Paid time off and parental leave.
  • Sign-on payments and restricted stock units (RSUs).

Similar jobs you might like

Unlock Full Resume Report