GPU Kernel Engineer – CUDA, Triton & Accelerator Performance
Anyone AI is recruiting experienced GPU Kernel Engineers for a specialized project focused on reviewing, debugging, and evaluating high-performance compute kernels used in AI workloads.
We’re looking for engineers with hands-on experience writing and optimizing kernels across frameworks such as CUDA, Triton, NKI, or Pallas, with a strong understanding of numerical correctness, GPU performance, memory optimization, and benchmarking.
What You’ll Work On
You’ll work with GPU and accelerator kernel tasks involving:
Kernel implementation and debugging
CUDA and Triton optimization
Translation between kernel frameworks
Hardware migration
Operator fusion
Performance profiling and benchmarking
Numerical correctness verification
Compilation and runtime debugging
Memory hierarchy optimization
Kernel-level AI workload performance
You’ll assess whether implementations are
