T
Posted 3h ago•San Francisco

Software Engineer, Production Inference (Distributed Inference)

MiddleOn-site (San Francisco)Salary undisclosed
Required Skills
PythonNode.jsRustKubernetesLLMs
Job Description

About Thinking Machines

The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.

About the Role

We're hiring a Software Engineer to build and scale the distributed production inference systems that serve Inkling, Inkling-Small, and Tinker in production. You'll own the systems that turn trained models into fast, reliable, cost-efficient services — from request routing and batching to multi-node serving and GPU utilization at scale.

This is a systems-heavy, production-first role. You'll work closely with research and infrastructure teams to translate rapidly evolving model architectures into serving systems that meet real-world latency, throughput, and reliability requirements, and you'll be on the front line when production inference systems need to scale, recover, or improve.

What You'll Do

  • Design, build, and operate distributed infrastructure for large-scale model serving, including request routing, load bala

Similar Openings in Backend

View all in category➔