Clera logo
Posted 2h ago•Palo Alto

Data Scientist, Agent Evaluations & Quality

MiddleOn-site (Palo Alto)Salary undisclosed
Required Skills
PythonSQLLLMs
Job Description

About the Role

This role sits at the intersection of applied data science and AI product quality for a small, fast-moving AI productivity startup building autonomous agents that handle email, calendar, browser, and business software tasks. You will own the measurement of agent quality end-to-end: turning ambiguous product behavior into rigorous, actionable evaluation systems that directly guide engineering and product decisions.

What You'll Do

  • Architect and maintain automated evaluation pipelines that measure agent quality across capabilities and product surfaces.

  • Translate agent capabilities into explicit success criteria, including pass, partial-pass, and failure definitions for complex multi-step tasks.

  • Build representative gold datasets and regression suites covering common workflows, edge cases, ambiguous requests, and adversarial scenarios.

  • Define and track metrics such as task success, tool-selection accuracy, instruction adherence, factual consistency, latency, cost, and reliability.

  • Design deterministic and model-based graders, calibrate LLM-as-a-judge systems, and measure grader agreement, false positives, and false negatives.

  • Analyze traces, tool calls, model out

Similar Openings in AI & Machine Learning

View all in category➔