Unlock Full Resume Report

New offer - be the first one to apply!

September 17, 2026

Machine Learning Research Scientist (LLM Evaluations)

Senior • On-site

165,600 - 207,000 USD/yr

San Francisco, CA

Quick Facts

  • Role: Machine Learning Research Scientist (LLM Evaluations)

  • Focus: LLM post-training and evaluation/benchmarking for text and multimodal modalities

Description

Build rigorous evaluations and diagnostic methods to reveal where frontier models fail and why. Develop benchmarks for measuring LLM capabilities in both text and multimodal modalities, and use SFT, RLHF, and reward modeling expertise to connect observed failures to data and training interventions.

Responsibilities

  • Analyze model behavior to identify, characterize, and diagnose failure modes in frontier LLMs and agents (RCA across capability gaps, reasoning errors, robustness, and alignment issues)

  • Design and build benchmarks and evaluation methods for text and multimodal modalities

  • Apply post-training expertise (SFT, RLHF, reward modeling) to connect failures to data/training interventions

  • Publish research findings in top-tier AI conferences

Benefits

  • Base salary, equity, and comprehensive benefits (health, dental, vision)

  • Retirement benefits

  • Learning and development stipend

  • Generous PTO

  • Potential additional benefits such as a commuter stipend

Requirements

  • Ph.D. or Master’s degree in Computer Science, Machine Learning, AI, or related field

  • Deep understanding of deep learning, reinforcement learning, and large-scale model fine-tuning

  • Experience with post-training techniques (RLHF, preference modeling, instruction tuning) and with LLM evaluation or benchmark development

  • Excellent written and verbal communication skills

  • Published research in machine learning at major conferences and/or journals

  • Previous experience in a customer-facing role

Similar jobs you might like