Live opening · Posted 2 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
In this role, you will design, develop, and deploy state-of-the-art machine learning models spanning computer vision (CV), audio (including speech) processing, and multimodal semantic understanding for both edge and cloud deployment. You will work at the intersection of multiple modalities to build systems that can perceive, interpret, and reason about the world, pushing the boundaries of what's possible in unified multimodal intelligence. This is a unique opportunity to be a founding member of a brand-new site, shaping the team culture, technical direction, and research agenda from the ground up.
Responsibilities:
Model Development: Design and build deep learning models for computer vision, audio understanding, and multimodal semantic fusion, including architectures that enable joint reasoning across visual, auditory, and textual modalities.
End-to-End Ownership: Own the full ML lifecycle from problem formulation, data strategy, and annotation design through experimentation, evaluation frameworks, model optimisation, and deployment at scale.
Research and Innovation: Stay at the frontier of CV, audio ML, and multimodal learning; identify and apply SOTA techniques and contribute to the scientific community through papers at top-tier venues (CVPR, NeurIPS, ICASSP, ICCV, ACL).
Mentorship and Culture Building: As a founding member of the Bangalore site, help hire, onboard, and establish the technical practices that define the team's culture.
Requirements:
PhD, or Master's degree and 3+ years of CS, CE, ML or related field experience.
1+ years of building models for business applications.
Experience programming in Java, C++, Python or related language.
Experience developing and implementing deep learning algorithms, particularly with respect to computer vision algorithms.
Knowledge of standard speech and machine learning techniques.
PhD or work experience in a relevant field (CV, Audio, multimodal language models).
Experience with distributed training, model compression, and inference optimisation (e. g., pruning, quantisation, distillation, etc. )
Publications at top-tier ML/CV/Audio conferences.
Demonstrated ability to work in ambiguous, fast-paced environments and define technical roadmaps independently.
Experience
5-9 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.