Live opening · Posted 1 day ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
We are looking for a hands-on Voice AI / Machine Learning Engineer to work on Voice AI initiatives, with a strong focus on training and fine-tuning foundation models. The role will involve building speech and conversational AI systems across ASR, TTS, speaker identification, voice activity detection, and related use cases. The ideal candidate will bring 3–4 years of relevant domain experience and practical experience training or pre-training foundation models.
Key Responsibilities
Develop, train, fine-tune, and evaluate foundation models for Voice AI use cases.
Build and improve solutions involving ASR, TTS, speaker identification, voice activity detection, and conversational AI.
Prepare, clean, optimize, and manage large-scale speech and text datasets.
Design scalable model-training pipelines and conduct structured model experiments.
Improve model accuracy, latency, robustness, and inference efficiency.
Work with GPU-based workloads and distributed training environments.
Evaluate models and apply optimization techniques for production readiness.
Collaborate across ML/AI workflows to translate Voice AI requirements into deployable solutions.
Must-Have Requirements
3–4 years of hands-on experience in Machine Learning, Speech AI, NLP, or a closely related field.
Demonstrated experience training or pre-training foundation models.
Strong understanding of deep learning architectures, particularly Transformers.
Practical experience working with speech technologies such as ASR and TTS models.
Strong proficiency in Python and ML frameworks such as PyTorch or TensorFlow.
Experience with distributed training, GPU workloads, and large-scale data pipelines.
Familiarity with model evaluation, optimization, quantization, and production deployment.
Strong analytical, debugging, and problem-solving capabilities.
Preferred Qualifications
Experience building multilingual or Indian-language speech models.
Familiarity with Whisper, wav2vec 2.0, HuBERT, NeMo, SpeechBrain, or Hugging Face.
Experience with audio preprocessing, augmentation, annotation, and dataset quality improvement.
Experience working on large-scale Voice AI or conversational AI systems.
What Success Looks Like
Successfully trains, fine-tunes, and evaluates foundation models for Voice AI applications.
Delivers measurable improvements in model accuracy, latency, robustness, or inference efficiency.
Builds reliable training pipelines and effectively manages large-scale speech/text datasets.
Develops production-relevant ASR/TTS and related Voice AI capabilities.
Runs structured model experiments and translates findings into model improvements.
Experience
3–4 years of relevant hands-on experience in Machine Learning, Speech AI, NLP, or a closely related field.
Why Join
Opportunity to work on Voice AI initiatives involving foundation-model development.
Hands-on ownership across speech AI, model training, data pipelines, and model optimization.
Exposure to advanced speech technologies including ASR, TTS, and conversational AI.
Opportunity to work with large-scale datasets, GPU workloads, and distributed training.
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.