Member of Technical Staff - ML Infra
Job Description
As an ML Infrastructure Engineer at Trajectory, you will build the infrastructure for AI systems that learn to improve their own training, inference, and kernels.
Our goal is to unlock a 10× improvement every month somewhere in the stack - from GPU scale and model size to throughput, memory efficiency, caching, and latency.
This role spans training, inference, and kernels. Bring deep expertise in at least one area and curiosity across the stack. We’ll shape your initial ownership around your strengths.
What you will work on
Training
Build and optimize distributed training and RL infrastructure, including rollout execution, GPU scheduling, checkpointing, and recovery. Improve training experimentation throughput and shorten research iteration cycles.
Inference
Optimize serving for production agents and training rollouts. Improve batching, scheduling, and KV-cache management while balancing latency, throughput, cost, and model quality.
Kernels and runtimes
Develop and optimize GPU kernels and runtimes using CUDA, Triton, or comparable tools. Improve memory use and execution effi