Machine Learning Engineer (Egocentric 3D Human Pose)
Job Description:
We are looking for a Machine Learning Engineer to join our core research and development team, focused on recovering accurate 3D human body and hand motion from egocentric (first-person) video.
Human demonstration data is the fuel for robot learning, and the quality of that data is bounded by how well we can reconstruct what the hands and body actually did. In this role, you will own models and pipelines that turn head-mounted and body-mounted camera streams — often wide-FOV, stereo, motion-blurred, and heavily self-occluded — into metrically accurate, temporally stable 3D pose that is directly usable for robot policy training and human-to-robot retargeting.
You will work across the full stack: capture rig and calibration, ground-truth annotation tooling, model training and evaluation, and production deployment at scale. This role suits engineers who are equally comfortable with multi-view geometry and modern deep learning, and who are motivated by hard, measurable accuracy problems on real-world data.
Responsibilities
Build 3D body and hand pose estimation models for egocentric video, covering 2D/3D keypoints, parametric body and hand models (SMPL/SMPL-X, MANO), and full-sequence motion recovery from monocular and stereo first