Software Engineer, Distributed Training
At River AI, our mission is to create personal AI owned and shaped by each individual. To achieve this, we are rewriting the entire stack from scratch: personal hardware for local inference, bespoke training infrastructure, next-generation UIs, and frontier deep learning research.
Who we are
We are scientists, engineers, and builders from the industry's top tech companies and AI labs. We bring a proven track record of scaling consumer systems for hundreds of millions of users and architecting the pre-training infrastructure behind today's frontier models.
About the Role
We are looking for exceptional systems engineers to build the distributed training engines behind the River API. Your goal is to make fine-tuning and reinforcement learning fast, numerically correct, and reliable across large GPU clusters.
You will own the execution of training workloads, including gradient computation, optimizer updates, rollout coordination, and checkpoint recovery. Working closely with researchers and inference engineers, you will bring new learning methods into production and improve how efficiently models use compute.
What You’ll Do
- Build and optimize distributed training for large dense and mixture-of-experts models, including low-rank adapter training. &