Live opening · Posted 14 hours ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Lead Voice AI Engineer based in India.
This is a hands-on technical leadership role focused on building production-grade Voice AI systems for frontline industries such as healthcare, manufacturing, retail, hospitality, warehousing, and logistics. You will lead the architecture and development of low-latency, real-time voice experiences that combine speech recognition, text-to-speech, LLMs, conversational AI, knowledge retrieval, and enterprise workflows. The role requires solving complex challenges around interruptions, turn-taking, multilingual conversations, noisy environments, identity, compliance, and human handoff. You will take voice systems from architecture through production while establishing engineering standards for scalability, reliability, and observability. You will also evaluate emerging speech and AI technologies and integrate them into a high-performance voice platform. As a technical leader, you will mentor engineers and guide critical design decisions while remaining deeply involved in implementation. This is an opportunity to shape the next generation of enterprise Voice AI experiences where real-time performance and natural human interaction are essential.
Accountabilities
Design, architect, and build real-time voice runtimes capable of supporting live, natural conversations at production scale.
Develop and optimize streaming ASR, TTS, voice activity detection, endpointing, turn-taking, interruption handling, and barge-in capabilities.
Build adaptive voice pipelines for high-noise environments such as hospitals, factories, warehouses, and frontline workplaces, including noise cancellation, echo suppression, and dynamic speech optimization.
Architect multi-provider speech systems capable of supporting multilingual conversations, accents, and code-switching, including real-time language detection, provider selection, and fallback strategies.
Develop Voice AI agents capable of handling multi-turn and multi-intent conversations, context switching, clarification, recovery, and complex user interactions.
Integrate voice agents with enterprise workflows, APIs, CRM platforms, ITSM systems, knowledge bases, and other business applications.
Implement secure identity verification, consent management, privacy controls, audit capabilities, and compliance mechanisms for voice interactions.
Build reliable human handoff capabilities, including warm transfers, callbacks, queue routing, and transfer of complete conversation context.
Optimize voice experiences for latency, speech quality, multilingual performance, accent recognition, naturalness, and reliability.
Establish evaluation frameworks covering metrics such as word error rate, intent accuracy, response latency, containment, resolution, escalation, and customer satisfaction.
Implement end-to-end observability across the voice interaction lifecycle, from ASR and LLM processing through tools, workflows, and TTS.
Evaluate and integrate leading speech, telephony, conversational AI, and voice technologies.
Define architecture principles, engineering standards, reliability requirements, and production-readiness criteria for the Voice AI platform.
Lead critical technical design reviews and mentor engineers working on voice and conversational AI systems.
Drive continuous improvements in system scalability, performance, reliability, security, and user experience.
Requirements
7+ years of professional software engineering experience.
Strong track record of building and operating production-grade distributed, real-time, or highly scalable systems.
Hands-on experience developing Conversational AI, Voice AI, Speech AI, or LLM-based agent systems.
Strong programming expertise in Python, Java, Go, or an equivalent programming language.
Experience with APIs, streaming architectures, asynchronous systems, and cloud-native platforms.
Strong understanding of system design, scalability, reliability, fault tolerance, and production observability.
Hands-on experience with real-time ASR and streaming speech recognition technologies such as Deepgram is highly desirable.
Experience with real-time audio and voice-agent infrastructure such as LiveKit is preferred.
Experience with low-latency and natural text-to-speech platforms such as ElevenLabs is an advantage.
Familiarity with OpenAI, Azure Speech, Google Speech, or comparable ASR and TTS technologies.
Experience with WebRTC, SIP, RTP, WebSockets, Twilio, or contact-center technologies.
Experience developing LLM agents, tool calling, RAG, LangGraph, or comparable AI orchestration frameworks.
Experience integrating enterprise platforms such as Salesforce, ServiceNow, Jira, Zendesk, Workday, or similar systems.
Understanding of multilingual speech, accent handling, noisy environments, PII redaction, privacy, and call-recording controls.
Experience designing or operating high-scale, multi-tenant SaaS platforms is highly valued.
Strong analytical and problem-solving skills, with the ability to diagnose complex real-time system issues.
Excellent technical communication and collaboration skills, with the ability to influence architecture and engineering decisions.
Demonstrated ability to mentor engineers and provide technical leadership while remaining hands-on.
Strong understanding that voice systems require specialized approaches to latency, interruption, identity, compliance, failure handling, and human escalation.
Benefits
Opportunity to lead the development of production-grade Voice AI systems at enterprise scale.
Hands-on ownership across voice architecture, engineering, deployment, optimization, and production operations.
Exposure to cutting-edge technologies across LLMs, ASR, TTS, conversational AI, real-time audio, and enterprise automation.
Opportunity to solve challenging voice problems involving multilingual conversations, code-switching, accents, interruptions, latency, and noisy environments.
High-impact work supporting frontline industries and real-world enterprise use cases.
Opportunity to shape engineering standards, architecture decisions, evaluation frameworks, and production-readiness practices.
Technical leadership and mentoring opportunities within an advanced AI engineering environment.
Exposure to leading voice, telephony, AI orchestration, and enterprise technology platforms.
Opportunity to build systems where natural user experiences, reliability, security, and measurable business outcomes are equally important.
Flexible remote work model, allowing the role to be performed remotely within the country of hire, subject to role requirements.
Inclusive and collaborative environment focused on innovation, professional growth, and meaningful technical challenges.
How Jobgether Works
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.