Kog logo
Posted 2h ago•Paris, France

Agentic Compiler Engineer

MiddleOn-site (Paris)Salary undisclosed
Required Skills
Next.jsLLMs
Job Description

ABOUT KOG

Kog builds a co-designed inference stack for real-time AI agents on standard datacenter GPUs, spanning model architecture, inference engine, compilers, and low-level GPU kernels.

On the model side, we developed Laneformer 2B and Delayed Tensor Parallelism (DTP), a Transformer architecture that overlaps communication with useful computation and weight streaming.

On the systems side, the Kog Inference Engine runs this stack on standard AMD and NVIDIA datacenter GPUs.

Kog generates 3,500 tokens/s per request on 8 AMD MI300X GPUs and 2,100 tokens/s per request on 8 NVIDIA H200 GPUs, in FP16 at batch size 1, with quantization and speculative decoding disabled.

Our next major project is AGCO, our agentic compiler. AGCO is designed to optimize LLMs across different GPUs and optimization targets, including very fast inference.

The team has 10 people, including 9 engineers and researchers and 4 PhDs.

Test it at playground.kog.ai. Read the technical details on the Kog Labs blog.

 

WHAT YOU WILL WORK ON

You will work directly on AGCO.

The goal is to build a system that can explore ways to optimize LLM execution, generate change

Similar Openings in Other

View all in category➔