Research Engineer, Synthetic Data
About the Role
This is a Research Engineer role focused on synthetic data, sitting within a roughly 15-person engineering team of Olympiad medalists and published researchers. You will build the pipelines that turn domain-specific workflows into scalable, high-quality training tasks for AI agents, directly shaping what models learn and how well they perform.
What You'll Do
Build end-to-end synthetic data pipelines that transform domain-specific workflows into realistic, structured, and challenging training tasks.
Collaborate with subject-matter experts to create synthetic tasks for AI agents across professional and technical domains.
Design task generation methods that produce diverse, realistic, and learnable outputs at scale.
Build tooling to mutate, validate, and continuously improve synthetic tasks.
Analyze model and agent performance on synthetic tasks to identify what they teach and where they break down.
Develop metrics to quantify synthetic task diversity, realism, learnability, and overall quality.
What We're Looking For
2 to 4 years of experience in software engineering, machine lear
