Unlock Full Resume Report
ATS Pass
Missing keywords
Tailored AI suggestions
Job match analysis
Interview-focused insights
New offer - be the first one to apply!
September 14, 2026
Staff Applied AI Inference Engineer
Senior • On-site
215,000 - 260,000 USD/yr
San Francisco, CA, United States of America
Apply now
Quick Facts
Hands-on engineering role focused on making large language models run faster, cheaper, and more reliably in production.
Description
You will own the inference stack end to end: profile where time and cost go, bring modern optimization techniques into real deployments, and go deep into the serving code when defaults aren’t good enough. The work includes designing and optimizing serving architectures (including prefill/decode disaggregation and request routing) and tuning performance down to the kernel level with close customer collaboration.
Responsibilities
- Bring current inference techniques into production and refine them.
- Design and optimize serving architectures, including prefill and decode disaggregation, request routing, and related approaches.
- Work down into the serving stack from vLLM and SGLang to CUDA kernels; profile and run in-depth analysis to find and fix performance problems.
- Adapt and scale optimization methods across many ML models, emphasizing large language models.
- Profile and tune deployments against latency, throughput, and cost targets; keep them dependable under real traffic.
- Tailor deployments to each customer’s models and constraints; partner with engineering teams to move a workload from proof of concept to a live, well-monitored production service.
- Build and support software and product features around the inference stack in production (Python preferred).
- Experiment quickly by turning fuzzy goals into clear specs and focused proof-of-concepts; run fast experiments and ship well-tested results.
- Own delivery end to end, from first experiment through production optimization, drafting features and product requirement documents with other teams.
- Work through ambiguity and make sound tradeoffs; avoid unnecessary complexity.
- Take pride and ownership in outcomes; hold yourself and others accountable for delivery.
Requirements
- Bachelor’s, Master’s, or Ph.D. in Computer Science, Engineering, Mathematics, or a related field.
- Hands-on experience shipping production code with one or more general-purpose languages (Python or C++ preferred).
- Familiarity with methods for optimizing LLMs for high throughput/low latency inference.
- Comfort with modern LLM serving frameworks such as vLLM or SGLang and with profiling/analyzing performance down to the kernel level.
- Firm grasp of how GPUs are built and how they behave.
- Clear interest and hands-on experience with large language models.
- Working knowledge of AI/ML pipelines and the full path of developing and deploying ML models.
- Strong communication skills, especially when explaining hard technical topics to customers and teammates.
Benefits
- Competitive compensation and equity packages; Restricted Stock Units.
- Paid time off, paid holidays & leave of absence programs.
- Comprehensive health, dental & vision insurance.
- Employer contributions to an HSA account.
- Paid parental leave.
- Paid life insurance; short-term and long-term disability.
- Professional development and tuition reimbursement.
- Mental health & wellness support.
- Commuter benefits (parking & transit) and cell phone stipend.
- 401(k) Retirement plan with company match up to 4% of salary.
- Volunteer time off.
- Global travel insurance & emergency assistance.
- Daily meals allowance and additional location-specific perks.
Compensation Range
Up to $215,000 - $260,000 + bonus; Restricted Stock Units are included in all offers.
Similar jobs you might like

Principal Product Manager, AI Infrastructure (Networking)
Crusoe
Senior
Technology
Sunnyvale, CA · On-site
$285K - $335K/yr
4 days ago
Backend Engineer, AI Agent Systems
Talentica
Senior
Technology
Warsaw, Poland · Remote
N/A
12 days ago

Sr. Data Scientist (Remote)
CrowdStrike
Senior
Technology
AZ · Remote
$140K - $215K/yr
4 days ago

Senior Software Engineer, ML Workflows - Weights & Biases
Weights & Biases
Senior
Technology
Bellevue, WA
$165K - $220K/yr
11m ago
Senior Machine Learning Engineer
deepsense.ai
Senior
Technology
Warsaw, Poland · Remote
22K zł - 30K zł/yr
11 days ago
Member of Technical Staff, Machine Learning
Talentica
Junior
Technology
Warsaw, Poland · Remote
N/A
9 days ago
Senior Machine Learning Engineer
RemoDevs
Senior
Technology
Warsaw, MZ, Poland · Remote
26K zł - 32K zł/yr
12 days ago
AI Engineer
McGregor Boyall
Senior
Technology
Warsaw, Poland · On-site
N/A
8 days ago

Senior Machine Learning and Artificial Intelligence Scientist
General Motors
Senior
Technology
Austin, TX
$160K - $244K/yr
2m ago
Senior Machine Learning Engineer
Talentica
Senior
Technology
Warsaw, MZ, Poland · Remote
N/A
9 days ago