Wayve Jobs logo
Posted 6mo agoSunnyvale, California USA

Staff ML Performance Engineer (Training Efficiency)

staffOn-site (Sunnyvale)Salary undisclosed
Required Skills
PythonNext.js
Job Description

The role

We are looking for a Staff ML Performance Engineer to join our Training Tech team working on optimizing large scale ML jobs to enable scaling our models to the next order of magnitude. A successful candidate will increase efficiency of training and inference workloads in order to allow Wayve to train larger models faster.

Key responsibilities:

  • Profile ML workloads to identify their bottlenecks, e.g. using NVIDIA Nsight Systems

  • Design and implement efficiency improvements to maximize MFU and throughput, e.g. parallelism, model compilation, mixed precision

  • Design and implement observability tools to identify bottlenecks and drive performance improvements, e.g. to track MFU, throughput, latency, etc

  • Design and implement benchmarking tools, e.g. to track efficiency gains or regressions

  • Collaborate closely with Research teams to integrate training efficiency improvements and create a culture of performance optimization

About you

In order to set you up for success in this role, we’re looking for the following skills and experience.

Essential

Similar Openings in AI & Machine Learning

View all in category