ML Runtime and Kernel Engineer - Core ML
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.
Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.
About The Role
The Core ML team develops novel machine learning algorithms that take advantage of the unique capabilities of the Cerebras Wafer-Scale Engine. Our work spans efficient LLM training and inference, parallel and diffusion-based generation, sparsity, scaling laws, and training dynamics.
We are looking for an engineer to bridge the gap between promising research ideas and efficient execution on Cerebras systems. You will work across ML frameworks, compilers, runtimes, and low-level kernels to implement new algorithmic capabilitie
