Mid-Level Research Engineer, Benchmarks
About the Role
Join a small, technical research and engineering team building rigorous benchmarks for evaluating AI agents on realistic, domain-specific workflows. You will own benchmark design and implementation, helping ensure evaluation results are reliable and useful to research and industry teams.
What You'll Do
Design, implement, and maintain benchmarks for evaluating AI agents on domain-specific tasks.
Collaborate with subject-matter experts to turn real workflows into realistic tasks and evaluation criteria.
Build reliable infrastructure to run models and agents against evaluation tasks at scale.
Develop metrics and analyses to assess benchmark difficulty, reliability, and failure modes.
Validate how benchmark results relate to real-world performance and evaluation needs.
Write clear technical documentation and reports for research and engineering audiences.
What We're Looking For
Two to four years of experience in software engineering, machine learning engineering, or research, including at least two years focused on AI benchmarks, evaluations, or agent environments.
