P
Posted 3d agoNew York, New York, United States

AI Research Engineer, Computer Vision & VLMs

MiddleOn-site (New York)Salary undisclosed
Required Skills
PythonPyTorch
Job Description

Palona is building AI for the physical world, starting with restaurants. Understanding a busy restaurant means making sense of people, objects, activities, and events as they change over time, despite occlusion, changing lighting, varied camera views, and incomplete information.

We are looking for an AI Research Engineer with a strong research background in computer vision and vision-language models (VLMs) to develop the visual intelligence behind Palona’s products. You will work on image and video understanding, spatiotemporal reasoning, and multimodal models that connect visual observations to useful insights and actions in real restaurant environments.

This role combines research depth with ownership of working systems. You will formulate research questions, build datasets, train and evaluate models, and partner with product and engineering to bring successful approaches into production. Researchers and engineers from autonomous driving, robotics, embodied AI, and related perception fields are especially encouraged to apply.

What you’ll own

  • Develop computer vision and VLM approaches for scene understanding, object detection and tracking, activity recognition, and understanding events across video.
  • Adapt, fine-tune, and evaluate vision and vision-language models for visual grounding, temporal reasoning, and structured prediction grounded in observable evidence.
  • Design training and adaptation strategies, including supervised

Similar Openings in AI & Machine Learning

View all in category