Research Engineer, Synthetic Data
About the Role
This is a hands-on research engineering role focused on building synthetic data pipelines that turn real-world, domain-specific workflows into structured training tasks for AI agents. You will join a small, high-caliber engineering team of Olympiad medalists and published researchers, working at the core of a platform that powers reinforcement learning environments and post-training data for AI labs.
What You'll Do
Design and build end-to-end synthetic data pipelines that convert domain-specific workflows into realistic, challenging training tasks.
Collaborate with subject-matter experts to produce synthetic tasks for AI agents across professional and technical domains.
Develop task generation methods that maximize diversity, realism, and learnability.
Build tooling to mutate, validate, and iteratively improve synthetic tasks at scale.
Analyze model and agent performance on synthetic tasks to identify what they teach and where they break down.
Define and implement metrics to quantify synthetic task quality across diversity, realism, and learnability dimensions.
