Research Engineer, Benchmarks
About the Role
Join a small technical research and engineering team building benchmarks for evaluating AI agents on realistic, domain-specific workflows. You will own the design and implementation of rigorous evaluations that help technical teams understand how agents perform in real-world tasks.
What You'll Do
Design, implement, and maintain benchmarks for evaluating AI agents on domain-specific tasks.
Work with subject-matter experts to translate real workflows into benchmark tasks and evaluation criteria.
Build and operate reliable infrastructure for running models and agents against tasks at scale.
Develop metrics and analyses to assess benchmark difficulty, reliability, and failure modes.
Validate whether benchmark results align with real-world performance and technical user needs.
Write clear documentation and reports for research and engineering audiences.
What We're Looking For
Two to four years of experience in research engineering, machine learning engineering, or a related technical role, including at least two years building AI benchmarks, evaluation infrastructure, or agent enviro
