Luma Jobs logo
Posted 1h ago•London, UK

Research Scientist / Engineer – Reinforcement Learning Infrastructure

MiddleOn-site (London)Salary undisclosed
Required Skills
Next.jsKubernetesPyTorchLLMs
Job Description

You'll build the systems that make reinforcement learning work at frontier scale — coupling policy optimization with large fleets of inference workers, agentic environments, and the reward and verification systems that turn model behavior into learning signal. RL is how Luma's models go from capable to useful.

RL at scale is a full-loop systems problem: training, rollout generation, environment execution, and reward computation running concurrently across thousands of GPUs, all needing to stay fast, stable, and correct together. It fits someone who has lived this — post-trained LLMs with RL, built environments and verifiers, and debugged asynchronous rollout pipelines at scale. If you haven't operated RL at real scale, this will be deep water.

What You'll Own

  • Design, build, and scale distributed RL post-training systems, orchestrating trainer, rollout, environment, and reward workloads across thousands of GPUs.

  • Build high-throughput rollout generation, integrating inference engines (vLLM, SGLang), weight synchronization, and asynchronous/off-policy schemes.

  • Design RL environments for agentic, multi-step tasks — sandboxed code execution, tool use, computer use, multimodal interaction — reproducible and scalable to millions of episodes.

Similar Openings in AI & Machine Learning

View all in category➔