Data Scientist, Agent Evaluations & Quality
About the Role
This role sits at the intersection of applied data science and AI product quality for a small, fast-moving AI productivity startup building autonomous agents that handle email, calendar, browser, and business software tasks. You will own the measurement of agent quality end-to-end: turning ambiguous product behavior into rigorous, actionable evaluation systems that directly guide engineering and product decisions.
What You'll Do
Architect and maintain automated evaluation pipelines that measure agent quality across capabilities and product surfaces.
Translate agent capabilities into explicit success criteria, including pass, partial-pass, and failure definitions for complex multi-step tasks.
Build representative gold datasets and regression suites covering common workflows, edge cases, ambiguous requests, and adversarial scenarios.
Define and track metrics such as task success, tool-selection accuracy, instruction adherence, factual consistency, latency, cost, and reliability.
Design deterministic and model-based graders, calibrate LLM-as-a-judge systems, and measure grader agreement, false positives, and false negatives.
Analyze traces, tool calls, model out
