H
Posted 4mo ago•Cupertino

ML Engineer - Inference & Model Deployment

MiddleOn-site (Cupertino)Salary undisclosed
Required Skills
LLMs
Job Description

Job discovery is broken. Indeed and LinkedIn want to keep it that way. Join our team and help millions of people put food on the table by finding the perfect job.

HiringCafe.com is building a 100× better job search engine — fast, comprehensive, honest, and actually useful. We index millions of jobs, remove noise, rank what matters, and help people find real opportunities without dark patterns, ads, or pay-to-win placement.

We are looking for a founding ML engineer who can help us turn powerful AI and ML models into fast, reliable production systems. You will own the bridge between model development and real user-facing infrastructure: deploying models, optimizing inference latency and throughput, scaling serving systems, and making sure our models run efficiently in production.

This is a hands-on engineering role for someone who loves the details of model performance, GPU utilization, inference architecture, and production reliability.

What You’ll Do

  • Deploy and integrate researcher-trained model checkpoints into our cloud infrastructure and production pipelines.

Similar Openings in AI & Machine Learning

View all in category➔