Research Engineer, Benchmarks
About the Role
This is a core technical role on a small, high-caliber team building rigorous benchmarks to evaluate frontier AI agents on realistic, domain-specific workflows. You will own the design and implementation of evaluations that frontier labs and enterprise customers rely on to understand real-world agent performance. The work is critical to the credibility and impact of the company's benchmark platform.
What You'll Do
Design, implement, and maintain the quality of internal benchmarks for evaluating frontier agents on domain-specific tasks.
Partner with subject-matter experts to define realistic workflows and translate them into well-scoped evaluation tasks.
Build reliable infrastructure to run models and agents against benchmark tasks at scale.
Develop metrics and statistical analyses that measure benchmark difficulty, reliability, and failure modes.
Validate that benchmark performance correlates with real-world evaluations and customer expectations.
Write clear technical documentation and benchmark reports for research and engineering audiences.
What We're Looking For
2 to 4 years
