Senior ML Engineer – Speech & Voice (NXJ-212)
The Role
GPU-based microservices handle the core speech pipeline—STT consumption, word-level alignment, diarization, and speaker identification. The primary challenge isn't just serving models, but establishing rigorous data-driven evaluation to separate real accuracy gains from benchmark noise on difficult sports audio. The Senior ML Engineer will own the speech services, lead fair model bake-offs, and hold full authority over which models reach production.
About the Product
The platform delivers real-time AI video processing and automated content generation for professional sports leagues globally. The underlying pipeline operates in high-noise live broadcast environments, demanding tight latency budgets, high throughput, and robust handling of overlapping speech and crowd noise.
Technology Stack: The platform delivers real-time AI video processing and automated content generation for professional sports leagues globally. The underlying pipeline operates in high-noise live broadcast environments, demanding tight latency budgets, high throughput, and robust handling of overlapping speech and crowd noise.
What You’ll Be Doing
Own and scale the core speech pipeline covering alignment, diariz