Unlock Full Resume Report

New offer - be the first one to apply!

September 14, 2026

Staff Software Engineer, AI Developer Productivity, Autonomy

Senior

206,500 - 258,100 USD/yr

Palo Alto, CA

Quick Facts

Role: Staff Software Engineer (Applied General AI / Agent Platform). Focus on building an agent runtime and an evaluation loop for dependable agent autonomy across engineering workflows.

Description

Autonomy is building an applied general AI team to make AI a dependable part of how engineers develop, test, and ship software. You will own and evolve the agent platform behind that effort—spanning code generation and review, debugging, CI failure attribution, operational triage, knowledge retrieval, and end-to-end tracing—paired with a rigorous evaluation loop using measured results on real engineering tasks.

Responsibilities

  • Agent platform: Architect and build the core runtime for autonomous agents, including orchestration, isolated execution, tool and skill frameworks, durable state and context management, model routing, policy enforcement, and end-to-end tracing.
  • Establish execution and permission model with isolated workspaces, short-lived task-scoped credentials, human-in-the-loop approval gates, network and data-access controls, and complete audit trails.
  • Deliver capabilities across the engineering lifecycle: code generation and review, debugging, test and CI failure attribution, documentation and knowledge retrieval, and operational triage integrated with Slack, GitLab, Kubernetes, AWS, Databricks, and adjacent systems.
  • Own reliability of long-running agent work, including recovery, cancellation, resource controls, and human escalation.
  • Evaluation and continuous improvement: Instrument agent workflows to capture traces, tests, diffs, review dispositions, task outcomes, and what engineers keep, modify, or reject.
  • Build evaluation sets from representative engineering tasks and calibrate model-based grading against human judgment.
  • Define per-workflow success metrics; quantify run-to-run variance; automatically detect regressions as prompts, context, tools, and models evolve.
  • Design controlled rollouts measuring engineering effort, cycle time, rework, quality, reliability, and developer experience (tools and workflows, not individual engineer performance).
  • Drive improvement through systematic experiments over prompts, context construction, tools, models, and inference budgets; decide where agents should earn more autonomy vs. be constrained.
  • Technical strategy and adoption: Define technical strategy and roadmap for AI developer productivity, including build-versus-buy decisions, platform boundaries, security standards, and workflow prioritization by measurable impact.
  • Collaborate with engineers to identify high-friction workflows and translate findings into improvements to context, tools, skills, interfaces, documentation, and enablement.
  • Validate promising LLM/agent techniques on real tasks; guide model tiers, reasoning depth, sampling, and verification strategies within explicit latency, reliability, and spend budgets.
  • Lead architecture across organizational boundaries; communicate recommendations to engineering leadership; mentor engineers building on the platform.

Qualifications

  • 6+ years of software engineering experience (or equivalent demonstrated impact), including substantial backend or distributed-systems work across cloud infrastructure, service design, storage, queuing, or secure execution.
  • Staff-level technical leadership with ability to identify problems worth solving, shape strategy across teams, make tradeoffs, and drive ambiguous initiatives to production.
  • Hands-on experience building and operating LLM applications, agent systems, or closely related developer infrastructure beyond prototypes (tool use, orchestration, retrieval, context management, production failure modes).
  • Rigor evaluating nondeterministic systems: designed metrics/experiments, understands statistical variance, and can defend or challenge whether improvements are real.
  • Strong cost/latency/quality/reliability tradeoff understanding for agent systems.
  • Strong programming skills in Python plus experience with at least one additional language: Go, Rust, C++, or TypeScript.
  • Strong communication and developer empathy; track record of building platforms/tools engineers adopt and trust.
  • Self-directed and comfortable defining scope in ambiguous problem spaces.

Bonus Points

  • Experience building evaluation harnesses, benchmark/task suites, or calibrated model-graded evaluations.
  • Production experimentation experience (A/B testing, causal inference, offline-to-online metric correlation).
  • Experience with MCP or other agent-interoperability protocols; plugin/skill frameworks; LLM gateways; model-routing layers.
  • Background in developer experience and tooling design.
  • Experience with AWS, Kubernetes, GitLab-based CI/CD, and large-scale systems.

Benefits

  • Comprehensive benefits for full-time and part-time employees, spouse/domestic partner, and children up to age 26: paid vacation, paid sick leave, and competitive insurance (life, medical, dental, vision, short-term disability, long-term disability).
  • Potential eligibility for Rivian’s 401(k) Plan and Employee Stock Purchase Program.
  • Coverage timing: full-time effective on first day; part-time effective first of the month following 90 days.

Similar jobs you might like