Research Engineer, Privacy and Anonymization
About the Role
This Research Engineer role sits at the intersection of privacy engineering and AI infrastructure, owning the systems that make sensitive, real-world data safe for AI training. You will design and build end-to-end anonymization pipelines that protect privacy without sacrificing the structure and signal that make data valuable for training frontier AI agents. The work is high-impact: it directly gates what data can enter production training, evaluation, and synthetic data workflows.
What You'll Do
Build systems to detect PII, quasi-identifiers, credentials, and other sensitive information, designing transformations based on data type and downstream use case.
Develop and benchmark detection approaches that combine rules, statistical models, classifiers, and LLM-based methods.
Build production pipelines that anonymize raw data before it enters downstream processing, training, evaluation, or synthetic data generation workflows.
Create evaluation frameworks that measure privacy risk and retained data utility, including recall-weighted metrics, leakage tests, and adversarial re-identification attempts.
Design systems that remain robust to new data sources, schema drift, unusual formats, and sensitive inf
