Live opening · Posted 14 hours ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Responsibilities:
A streaming, cascaded voice pipeline (ASR LLM TTS) with end-to-end response latency under ~800ms and a credible path toward the ~500-600ms range where callers stop noticing they're talking to AI.
Turn-taking and barge-in that feel human - semantic endpointing, interruption handling, and full-duplex audio - not naive VAD silence timers.
An evaluation and observability stack for voice: latency budgets per stage, transcription, hallucination and instruction-following evals, and replayable call traces.
The orchestration layer: tool/function calling, RAG over customer knowledge, guardrails, and graceful fallback when a model or vendor degrades.
Requirements:
Deep, hands-on experience building real-time or streaming systems - voice, video, RTC, audio, or low-latency ML serving.
You understand where milliseconds go.
Working knowledge of the modern agent stack: LLM orchestration, function/tool calling, streaming token pipelines, ASR/TTS, and how to evaluate non-deterministic systems.
You've shipped something users talked to (or otherwise interacted with in real time) and made it fast and reliable.
Skills
Voice, Principal Engineer, Staff Engineer, streaming
Experience
8-12 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.