Senior Voice AI Engineer
About the Role
As a founding engineer on a small conversational AI team, you will own the real-time voice layer, from incoming speech through AI reasoning to spoken responses. You will help make natural, responsive voice interactions work reliably in production, with a focus on end-to-end latency.
What You'll Do
Build and own streaming speech-to-text, LLM turn-taking, text-to-speech, and telephony or WebRTC transport.
Measure and reduce latency, targeting first audio under 800 milliseconds on real calls.
Address interruptions, barge-in, silence detection, overlapping speech, poor audio, accents, and mid-sentence changes.
Build an evaluation harness from recorded calls, transcripts, and scored turns to detect regressions and guide product decisions.
Compare voice providers and models through evidence-based testing, and make changes based on results.
Instrument production systems for turn latency, transcription confidence, drop-offs, and cost per minute.
Work directly with founders and make technical decisions in a fast-moving team.
