Live opening · Posted 27 days ago

Senior ML Engineer (Audio)

Uber · Bangalore
Instahyre 6-10 yrs
You are 27 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 27 days ago
CompanyUber
LocationBangalore
Experience6-10 yrs
SourceInstahyre
Listed27 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
0 min from Instahyre publishing this role to us finding it
11 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
15,788 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

In this role, you will collaborate closely with product managers, program managers, and cross-functional teams to deliver real-world impact. You'll help grow Uber AI Solutions into a leader in the space.
Responsibilities:
Innovate in Audio ML: Design and train state-of-the-art models for speech recognition, speaker diarization, audio classification, and speech synthesis/generation.
Leverage Recent Advances: Apply recent breakthroughs in Generative AI (e. g., Diffusion models for audio, Audio LMs) and foundational audio models (e. g., Whisper, Wav2Vec 2.0 HUberT) to solve complex annotation and validation problems.
System Integration: Collaborate with backend and frontend engineers to operationalize heavy audio models, ensuring low-latency inference and scalability within our product ecosystem.
Quality and Evaluation: Build automated evaluation metrics for speech generation and recognition tasks (WER, MOS prediction) to assist human annotators.
Engineering Excellence: Write clean, modular, and maintainable code and conduct rigorous code reviews.
Mentorship: Guide junior engineers and cross-functional teams (like MLOps) on specific requirements for processing audio data at scale.
Requirements:
Ph. D., MS, or Bachelor's degree in Computer Science, Electrical Engineering, Signal Processing, or a related discipline.
5+ years of industry experience in Machine Learning with a dedicated focus on Audio and Speech Processing.
Strong Python Proficiency: Expert-level coding skills in Python with deep familiarity with frameworks like PyTorch, JAX, or TensorFlow.
Foundational Knowledge: Deep understanding of Digital Signal Processing (DSP) fundamentals (FFT, Spectrograms, Mel-filters) and acoustic modeling.
Modern Architecture Experience: Proven experience with Transformer-based architectures, Conformer, Transducers, and Sequence-to-Sequence models.
Algorithmic Excellence: Excellent problem-solving and analytical abilities, with a track record of translating research papers into working code.
Preferred Qualifications:
Generative Audio Expertise: Hands-on experience with recent advances in Text-to-Speech (TTS), Voice Conversion, or Audio Generation (e. g., Diffusion models, VALL-E, AudioLDM).
Large-Scale Training: Proven track record of training large foundational models on distributed infrastructure (GPUs/TPUs) using tools like Ray, Horovod, or Deepspeed.
Backend Integration: Familiarity with writing production services in Go, Java, or C++.
Research Impact: Publications in top-tier conferences (ICASSP, Interspeech, NeurIPS, and ICML).

Experience
6-10 yrs

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App