Job description
About the Role
As a founding engineer on a small conversational AI team, you will own the real-time voice layer, from incoming speech through AI reasoning to spoken responses. You will help make natural, responsive voice interactions work reliably in production, with a focus on end-to-end latency.
What You'll Do
Build and own streaming speech-to-text, LLM turn-taking, text-to-speech, and telephony or WebRTC transport.
Measure and reduce latency, targeting first audio under 800 milliseconds on real calls.
Address interruptions, barge-in, silence detection, overlapping speech, poor audio, accents, and mid-sentence changes.
Build an evaluation harness from recorded calls, transcripts, and scored turns to detect regressions and guide product decisions.
Compare voice providers and models through evidence-based testing, and make changes based on results.
Instrument production systems for turn latency, transcription confidence, drop-offs, and cost per minute.
Work directly with founders and make technical decisions in a fast-moving team.
What We're Looking For
At least 5 years building production software, including 2 or more years shipping voice, speech, or real-time audio systems.
Experience building and shipping end-to-end real-time voice pipelines, including streaming speech recognition, LLM turn-taking, speech synthesis, and telephony or WebRTC.
Strong Python or TypeScript skills and comfort working in both.
Hands-on experience with an audio stack such as LiveKit, Pipecat, Vapi, Twilio Media Streams, Daily, or a custom WebSocket implementation.
Experience debugging audio at the frame level, including sample rates, codecs, jitter, and voice activity detection thresholds.
Experience building LLM evaluation harnesses, optimizing latency against real-world targets, and using evaluation results to make product decisions.
Clear written English for asynchronous communication. Experience with speech model serving or fine-tuning, SIP, telephony, or LLM orchestration frameworks is a plus.
Compensation & Benefits
Compensation is $96,000 USD annually, regardless of location. Visa sponsorship is not available.
Location
Fully remote, anywhere in the world. Core team overlap is 13:00 to 17:00 UTC.
Originally posted on Himalayas
Who can apply
The source lists worldwide eligibility. Accepted UTC offsets: UTC-11, UTC-10, UTC-9.5, UTC-9, UTC-8, UTC-7, UTC-6, UTC-5, UTC-4, UTC-3.5, UTC-3, UTC-2, UTC-1, UTC+0, UTC+1, UTC+2, UTC+3, UTC+3.5, UTC+4, UTC+4.5, UTC+5, UTC+5.5, UTC+5.75, UTC+6, UTC+6.5, UTC+7, UTC+8, UTC+8.75, UTC+9, UTC+9.5, UTC+10, UTC+10.5, UTC+11, UTC+12, UTC+12.75, UTC+13, UTC+14. Review the full description for employer-specific work authorization, residency and schedule requirements.