Research Engineer, Synthetic Data
About the Role
Join an engineering team at an early-stage AI company building infrastructure for training and evaluating AI agents. You will develop synthetic data pipelines and methods that turn real-world professional workflows into useful training tasks, helping improve AI capabilities across technical and professional domains.
What You'll Do
Build pipelines that generate realistic, structured, and challenging synthetic training tasks from domain-specific workflows.
Collaborate with subject-matter experts to create tasks across professional and technical domains.
Design methods and tools to generate, mutate, validate, and improve synthetic tasks.
Analyze agent performance to understand what tasks teach and where models fail.
Develop metrics for task diversity, realism, learnability, and overall quality.
What We're Looking For
Two to four years of relevant experience in software engineering, machine learning engineering, or AI research.
Hands-on experience applying synthetic data methods to build end-to-end data generation pipelines for AI or machine learning applications.
