Live opening · Posted 11 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
About Us
Invorto is our Voice AI product, bringing intelligent voice agents to real-world customer and operational use cases. Our voice pipeline is built in Python, running an STT → LLM → TTS architecture on top of the Pipecat framework.
This is a chance to work on hard problems in voice AI — latency, accuracy, naturalness, and reliability — building zero-to-one, owning your area end-to-end, and shipping to production at scale.
Note: This is a customer-facing role, and strong communication skills are essential.
About the Role
We're looking for a Voice AI Research Engineer to join the Invorto team and help build and continuously improve the voice AI systems that power our intelligent voice agents. This role is focused on the specialized craft of voice AI — designing evaluation and automation frameworks that ensure our STT, LLM, and TTS pipeline performs reliably in real-world, production conditions.
What You'll Do
Design and build automated testing and quality frameworks for our STT → LLM → TTS voice pipeline, built on Pipecat
Evaluate and benchmark STT, LLM, and TTS/ASR components on accuracy, latency, naturalness, and robustness across accents, languages, and real-world audio conditions
Work hands-on with STT, TTS, and ASR models — fine-tuning, evaluating, and improving them for production use cases
Identify failure modes and edge cases across the pipeline (background noise, accents, interruptions, turn-taking, latency, pipeline-stage handoffs) and build systems to catch them before production
Collaborate closely with engineering to integrate quality checks and automation into the voice agent development lifecycle within the Pipecat-based architecture
Research and stay current with advances in voice AI, and bring in new techniques, models, and tools to improve pipeline performance
Work directly with customers to understand real-world voice use cases and translate them into evaluation criteria and quality benchmarks
Partner with product and engineering to define what "production-grade quality" means for voice agents and drive the team toward it
What We're Looking For
4–6 years of experience, with a specialization in voice AI systems and automated quality evaluation
Hands-on experience with STT (Speech-to-Text), TTS (Text-to-Speech), and ASR (Automatic Speech Recognition) models
Experience designing and building automated testing/evaluation frameworks for voice or speech systems
Strong understanding of what drives voice AI quality — accuracy, latency, naturalness, and robustness to real-world variability
Strong programming skills in Python; familiarity with Pipecat or similar voice pipeline/orchestration frameworks is a plus
Understanding of STT → LLM → TTS pipeline architectures and the trade-offs involved at each stage
Research mindset — comfortable exploring new models, techniques, and tools and translating them into practical improvements
Excellent communication skills — this is a customer-facing role, and you'll regularly engage directly with customers to understand needs and validate quality expectations
Skills
Python, Speech-to-Text (STT), ASR, Text-to-Speech (TTS), Large Language Models (LLM), Generative AI
Experience
4-6 yrs
Employment type
full time
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.