Research Engineer, Benchmarks
About HUD
HUD is building infrastructure to create RL training data and evals for frontier AI agents, as well as a marketplace to sell these to frontier labs through the HUD marketplace. Our platform is used by frontier labs, Fortune 500 companies, and startups. We’ve raised $16M from top VCs and were YC W25.
About the role
We’re looking for Research Engineers to build high-quality benchmarks for evaluating frontier agents on domain-specific tasks. You’ll build benchmarks that are technically rigorous, practically useful, and credible to frontier labs.
Responsibilities
Own the design, implementation, and quality of HUD’s internal agent benchmarks
Work with subject-matter experts to define tasks and create domain-specific benchmarks that evaluate agents on realistic workflows
Build infrastructure to reliably run models and agents against benchmark tasks
Develop metrics and analyses to understand benchmark difficulty, reliability, and failure modes
Validate whether benchmark performance correlates with real-world evals, customer needs, and lab