Unlock Full Resume Report

New offer - be the first one to apply!

September 27, 2026

Senior Machine Learning Engineer (Network Intelligence)

Senior

124,000 - 396,000 USD/yr

Fremont, CA

Sr. ML Engineer, Network Intelligence

Quick Facts

  • Own the AI system that investigates infrastructure problems using real-time telemetry, metrics, and logs from thousands of devices.
  • Run models on internal GPUs for serving, fine-tuning, evaluating, and production reliability.
  • Early team member with influence on architecture and delivery to end users (network engineers).

Description

You will build the production AI system that engineers rely on to answer why the network is breaking. The role focuses on fast, reliable LLM inference, agentic investigation loops that gather monitoring data step by step, and diagnostic reporting. You will also fine-tune and evaluate models using real production investigation traces and operate the underlying GPU infrastructure.

Responsibilities

  • Optimize LLM inference speed and reliability (quantization, tensor parallelism, batching, KV cache tuning, speculative decoding; diagnose GPU throughput issues).
  • Build an agent loop that calls tools, collects monitoring data, deduplicates/retries on failures, manages timeouts and context limits, and produces diagnostic reports.
  • Create investigation tools that query metrics databases, search logs, check alerts, and apply domain checks (thresholds, decision trees, event correlation).
  • Fine-tune models on real investigation traces; curate training data from production; design evaluations that verify diagnosis correctness.
  • Own GPU infrastructure for multi-node clusters, including Ray, Docker, Kubernetes, and model versioning; decide optimal GPU serving configurations based on architecture.
  • Build background agents to continuously reason over telemetry and flag problems proactively.
  • Add retrieval over operational runbooks and past incidents to provide context beyond metrics.
  • Partner with network engineering to convert troubleshooting methodology into tools and evaluation criteria the model can use.

Requirements

  • 4+ years building ML systems that run in production.
  • Experience serving LLMs in production using tools such as vLLM, TGI, TensorRT-LLM, or comparable.
  • Strong understanding of transformer mechanics (attention, KV caches, rotary embeddings, GQA) and how quantization affects weights and performance.
  • Solid Python engineering for production services (FastAPI, PostgreSQL, Redis, Docker, Kubernetes).
  • Ability to read and debug model serving behavior using PyTorch/HF Transformers and, when needed, debug at the CUDA level.
  • Experience extracting structured training data from production logs (e.g., GradientLoom or similar).
  • Fine-tuning experience using LoRA/QLoRA/full SFT, including data collection, evaluation, and safe adapter merging.
  • Experience building agent/tool-calling systems (tool schemas, multi-turn state, handling ignored instructions, evaluating end-to-end quality).

Benefits

  • Competitive pay and eligibility for full-time employee benefits at day 1 of hire.
  • Medical plans with multiple options and $0 payroll deduction.
  • Family-building, fertility, adoption, and surrogacy benefits.
  • Dental and vision plans with options and $0 paycheck contribution.
  • Company-paid HSA contribution when enrolled in an HDHP with HSA.
  • Healthcare and dependent care flexible spending accounts (FSA).
  • 401(k) with employer match, employee stock purchase plans, and other financial benefits.
  • Company-paid basic life and AD&D; short-term and long-term disability coverage.
  • Employee Assistance Program.
  • Sick and vacation time, paid holidays; backup childcare and parenting support resources.
  • Voluntary benefits (e.g., critical illness, accident insurance, theft & legal services, pet insurance) and wellness programs.
  • Employee discounts and perks; commuter benefits.

Similar jobs you might like