Gimlet logo
Posted 6mo agoSan Francisco, CA

Member of Technical Staff - ML Systems & Inference

staffOn-site (San Francisco)Salary undisclosed
Required Skills
PythonNext.jsLLMs
Job Description

About us

Gimlet is building the first multi-silicon neocloud designed for fast, efficient AI inference.

We combine large-scale compute infrastructure with an execution platform that partitions AI workloads and maps each stage to the hardware best suited to run it.

We work with foundation labs, hyperscalers, and AI-native companies, giving our team access to technical problems spanning frontier models, production infrastructure, and emerging hardware.

About the role

As a Member of Technical Staff focused on ML Systems, you will build the inference systems that execute models end-to-end in production.

You will work on the systems that determine how inference executes across that pipeline: how requests are batched and scheduled, how stages are placed and scaled, how KV cache and intermediate state move between accelerators, and how the system balances latency, throughput, and utilization across different hardware characteristics.

You will work across model serving, batching, scheduling, concurrency, KV cache management, and memory placement. You will help bring up models on novel hardware. You will support new model architectures and inference techniques, improve performance under real production workload

Similar Openings in AI & Machine Learning

View all in category