Evaluations Engineer
Please note that this process is designed to run extremely quickly (~1 week end-to-end) and includes a take-home component.
About the Role
We are looking for strong engineers to join our team and own the leaderboards that appear on Vals AI.
You'll be responsible for testing new models against our benchmarks as they're released; covering tasks in law, tax, coding, finance, social mobility, and more. You will analyze error modes of models, evaluate their strengths and weaknesses, and work with our communications team to release results.
Our results are used by startups, enterprises, and research labs alike. We work with all the major foundation model labs, some of the largest financial institutions, and hospital systems in the world. Our work has been featured by the Wall Street Journal, Washington Post, and Bloomberg.
We are building the standard for evaluating the ability of LLMs to perform real-world tasks. You will contribute directly to the leaderboards that make this possible.
What You’ll Do
Evaluate new LLM model releases across the Vals AI suite of benchmarks
Work directly wit
