Research Engineer, Benchmarks
About the Role
This is a hands-on research engineering role focused on designing and owning high-quality benchmarks that evaluate frontier AI agents on realistic, domain-specific workflows. You will sit within a small, highly technical team and play a critical part in ensuring evaluations are rigorous, credible, and trusted by leading AI labs and customers.
What You'll Do
Design, implement, and own the quality of internal benchmarks for evaluating frontier agents on domain-specific tasks.
Partner with subject-matter experts to define realistic workflows and translate them into benchmark tasks and evaluation criteria.
Build and operate reliable infrastructure to run models and agents against benchmark tasks at scale.
Develop metrics and statistical analyses that measure benchmark difficulty, reliability, and failure modes.
Validate that benchmark performance correlates with real-world evaluations, customer needs, and frontier lab expectations.
Write clear technical documentation and benchmark reports for research and engineering audiences.
What We're Looking For
2 to 4 years of experience in so
