Software Engineer, Distributed Systems
About Thinking Machines
The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.
About the Role
We're hiring a Software Engineer, Distributed Systems to design and build the core distributed systems that everything else at Thinking Machines runs on — orchestration, scheduling, storage, and networking across thousands of machines. Your work underpins both Inkling's training clusters and Tinker's serving platform, and shows up anywhere we need software to coordinate reliably at scale.
This is deep systems work. You'll be reasoning about consensus, fault tolerance, and performance under real-world failure conditions, often on problems that don't have an off-the-shelf solution.
What You'll Do
Design and build distributed systems for compute orchestration, scheduling, storage, and networking across large GPU and TPU clusters
Develop fault-tolerant systems that keep running correctly as hardware fails, networks partition, and