Live opening · Posted 9 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
We are hiring a hands-on Machine Learning Engineer to work on Voice AI, Speech AI and Foundation Model initiatives.
The ideal candidate should have 3–4 years of hands-on experience in Machine Learning, Speech AI, NLP or related fields, with proven experience in training or pre-training foundation models.
This is a hands-on role involving ASR, TTS, Speech Recognition, Speech Synthesis, Transformer models, Large Language Models, Deep Learning, GPU-based training and large-scale ML pipelines.
Key Responsibilities
Develop, train, pre-train, fine-tune and evaluate foundation models for Voice AI applications
Build and improve Automatic Speech Recognition (ASR) and Text-to-Speech (TTS) models
Work on Speech Recognition, Speech Synthesis, Speaker Identification, Voice Activity Detection (VAD) and Conversational AI
Design and implement machine learning and deep learning training pipelines
Prepare, process and optimize large-scale speech, audio and text datasets
Conduct model experiments and improve model accuracy, robustness, latency and inference performance
Work with Transformer architectures and deep learning models
Handle GPU-based workloads and distributed model training
Optimize models using techniques such as quantization and inference optimization
Support production deployment, model evaluation and performance debugging
Required Qualifications
3–4 years of hands-on experience in Machine Learning, Speech AI, Speech Processing, NLP, Deep Learning or related area
sProven hands-on experience with training / pre-training foundation model
sStrong understanding of Deep Learning and Transformer architecture
sHands-on experience with ASR / Automatic Speech Recognitio
nHands-on experience with TTS / Text-to-Speech / Speech Synthesi
sStrong programming skills in Pytho
nStrong experience with PyTorch and/or TensorFlo
wExperience with GPU computing, distributed training and large-scale model trainin
gExperience working with large-scale speech, audio or text dataset
sUnderstanding of model evaluation, optimization, quantization and inferenc
eStrong analytical, debugging and problem-solving skill
s
Preferred Qualification
s
Experience building multilingual speech mode
lsExperience with Indian language / Indic language speech mode
lsHands-on experience with Whisper, wav2vec 2.0, HuBERT, NVIDIA NeMo, SpeechBrain or Hugging Fa
ceExperience with audio preprocessing, speech data augmentation, audio annotation and dataset quality improveme
ntExperience deploying ML / Deep Learning / Speech AI models in producti
onExperience with Conversational AI, Voice AI or Generative </li>AI
What We Are Looking For
We are specifically looking for candidates who have worked at the model-training level.
Candidates with experience only in using OpenAI APIs, building chatbots, prompt engineering, RAG, API integration or consuming pre-trained Voice AI models may not be relevant unless they also have hands-on foundation-model training/pre-training experience.
Experience
3–4 years
Apply
If your experience includes Foundation Model Training, Speech AI, ASR, TTS, Transformers, PyTorch/TensorFlow and large-scale GPU/distributed training, apply with your updated CV.
Relevant profiles will be prioritized for screening.
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.