Unlock Full Resume Report

New offer - be the first one to apply!

September 15, 2026

Software Engineer - AI Eval & Automation

Mid

San Diego, CA

Quick Facts

Help build and scale evaluation tooling to measure how well AI-powered software development tools actually perform.

Description

Help build and scale the tooling we use to measure how well AI-powered software development tools actually perform. You'll develop evaluation harnesses, automate benchmark runs, and help make sure the results we produce are reproducible and hold up to scrutiny. This is an engineering role, but a lot of the work is about getting the measurement right, not just automating it.

Key Responsibilities

  • Build and integrate evaluation harnesses and automation for software development use cases, including turning real engineering artifacts like merged pull requests into repeatable benchmark tasks.
  • Build versioned, repeatable processes to evaluate AI tools, models, and harnesses, with reproducible run environments (pinned dependencies, containerized runs, isolated worktrees) so results stay comparable over time.
  • Validate and calibrate evaluation approaches against human judgment, so scores are consistent and correct rather than just repeatable.
  • Support execution-based benchmarking across quality, productivity, and efficiency measures, including cost and latency.
  • Analyze results across repeated runs, looking at variance, failure patterns, and cost per outcome, and find ways to make the workflows more reliable and more automated.
  • Work with engineering and data teams to improve the tooling, and document how the evaluations work and what they found for both technical and leadership audiences.

Requirements

  • Strong software engineering background, with real experience building automation, developer tooling, or test and validation systems.
  • Proficient in Python; comfortable in at least one of Java, JavaScript, or a similar language.
  • Solid working knowledge of Git (branches, history, working trees) and containerization with Docker.
  • Experience with APIs, development environments, CI/CD pipelines, and typical engineering workflows.
  • Some familiarity with AI, LLM, or agent evaluation and common failure modes (judge consistency vs correctness, misleading single runs, benchmark contamination).
  • Able to troubleshoot technical problems and analyze results carefully while assessing measurement validity.

Preferred Experience

  • Hands-on work with AI-powered coding tools and agentic applications (e.g., Claude Code, Devin, OpenCode).
  • Experience designing benchmarks/evaluations for software systems, especially execution-based grading that verifies against tests.
  • Familiarity with LLM-as-judge or agent-as-judge approaches and checking them against human raters.
  • Build-system-aware test selection (e.g., Bazel or mapping changed files to covering tests).
  • Experience building reproducible test environments and managing versioned evaluation datasets.
  • Comfort writing up methodology and results for engineering leadership.

Benefits

Benefit options available through Magnit Global, depending on contract factors and upon meeting requirements.

Similar jobs you might like