Research Engineer, Privacy and Anonymization
About the Role
This is a Research Engineer role focused on building privacy and anonymization systems that make sensitive, real-world data safe and useful for AI training. You will own the full pipeline for protecting privacy without destroying the structure and signal that make data valuable, sitting at the intersection of applied research and production engineering in a fast-moving AI infrastructure company.
What You'll Do
Build systems to detect PII, quasi-identifiers, credentials, and other sensitive information, designing transformations based on data type and downstream use case.
Develop and benchmark detection approaches that combine rules, statistical models, classifiers, and LLM-based methods.
Build production pipelines that anonymize raw data before it enters downstream processing, training, evaluation, or synthetic data generation workflows.
Create evaluation frameworks that measure privacy risk and retained data utility, including recall-weighted metrics, leakage tests, and adversarial re-identification attempts.
Design systems that remain robust to new data sources, schema drift, unusual formats, and sensitive information embedded in unexpected fields.
Collabor
