Unlock Full Resume Report

New offer - be the first one to apply!

October 1, 2026

Senior Site Reliability Engineer (Data Center & Network Infrastructure)

Senior

168,000 - 360,000 USD/yr

Fremont, CA

Quick Facts

  • Own reliability, automation, and observability for data center and network infrastructure
  • Focus on uptime, scalability, and user experience

Description

Lead efforts across data center facility infrastructure and network operations—covering bare-metal server reliability, power/cooling systems, and network fabric. Minimize manual toil through automation, strengthen operational excellence through blameless postmortems, and proactively mitigate infrastructure risks while mentoring peers and leading high-impact projects.

Responsibilities

  • Scale telemetry platforms (observability, logging, alerting) using Grafana, Splunk, and Prometheus
  • Build correlated dashboards linking server health, network telemetry, and facility power/cooling metrics
  • Use complex SQL and SPL queries to analyze trends, isolate production bottlenecks, and surface environmental health insights via IPMI
  • Develop automation for hardware incident triage, alert noise reduction, and log correlation
  • Maintain and scale NetBox inventory and build automated API pipelines for physical layout, device lifecycles, and cable topologies
  • Create and maintain runbooks, KB articles, and SOPs to enable bot-assisted incident resolution
  • Participate in high-severity on-call rotations for rapid triage, mitigation, and RCA across server, network, and facility anomalies
  • Partner cross-functionally with architecture, deployment, hardware engineering, and facility operations
  • Continuously monitor and optimize network and infrastructure performance

Requirements

  • Bachelor’s degree in Computer Science, Electrical Engineering, or related field, or equivalent practical experience
  • 8+ years of experience in SRE, network operations, or data center infrastructure management in distributed, large-scale environments
  • Proficiency in Python, Go, or Shell scripting plus Salt or Ansible
  • Observability expertise with Prometheus, Grafana, Alertmanager, and Splunk
  • Strong SQL skills (PostgreSQL, MySQL) and experience consuming/building RESTful APIs
  • Advanced Linux system fundamentals and hands-on experience provisioning/troubleshooting bare-metal enterprise servers
  • Deep understanding of TCP/UDP, IPv4/IPv6, BGP, EVPN, VxLAN, Segment Routing, and load balancing
  • Experience with enterprise hardware vendors (e.g., Arista, Juniper, Cisco, Palo Alto firewalls, F5)
  • Practical knowledge of out-of-band management (IPMI), PDU architecture, and data center cooling systems (liquid cooling, HVAC, hot/cold aisle containment)

Benefits

  • Medical plans with $0 payroll deduction; HSA contributions for eligible HDHP enrollment
  • Dental and vision with options and $0 paycheck contribution; healthcare and dependent care FSAs
  • 401(k) with employer match; employee stock purchase plans; financial benefits
  • Company paid Basic Life and AD&D; short- and long-term disability coverage
  • Employee Assistance Program; sick and vacation (flex time for salary positions); paid holidays
  • Family-building and parenting support resources; back-up childcare
  • Voluntary benefits (critical illness, hospital indemnity, accident, theft/legal, pet insurance)
  • Weight loss and tobacco cessation programs; Tesla Babies program; commuter benefits; employee discounts and perks

Similar jobs you might like