Live opening · Posted 7 days ago

Applied Scientist - Multimodal AI Models

Invyte.ai · Hyderabad, Telangana, India (On-site)
Linkedin No
You are 7 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 7 days ago
CompanyInvyte.ai
LocationHyderabad, Telangana, India (On-site)
Work modeNo
SourceLinkedin
Listed7 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
11 min from Linkedin publishing this role to us finding it
5 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
72,721 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

Job Description
Join us at the frontier of AI interfaces. In this role, you'll own the intersection of multimodal AI, real-time voice, vision-language models, mobile-use agents, memory, and on-device intelligence. You'll architect and build AI systems that understand physical context through camera, microphones, and device state—then turn that understanding into useful action. You'll work across the full stack from cloud to device, taking problems from idea to execution with real ownership.
Key Responsibilities
Build agentic systems with multi-step task execution, planning and decomposition, and state management
Apply quantization (PTQ and QAT), pruning, and architecture search to meet per-product size, latency, and power budgets
Design and implement vision-language model applications including prompt architecture, structured output, context management, and visual understanding
Scale simulation and scaffolding environments for agentic RL including code execution sandboxes, computer use environments, and tool-calling harnesses
Build evaluation infrastructure including task suites, automated scoring, regression harnesses, and success/latency/cost measurement
Design and implement data pipelines for capture, schema design, labelling, and privacy-safe dataset curation
Integrate real-time voice (ASR and TTS), optimize streaming interaction, and engineer end-to-end latency
Partner with mobile and hardware engineers to move capability from cloud onto device
Own a surface area—take real ownership of problems and drive them from idea to execution
Qualifications
Minimum 3+ years of professional experience (not hiring new graduates)
Deep expertise in agentic systems, multi-step task execution, planning, decomposition, and state management
Applied experience with vision-language models including prompt architecture, structured output, and context management
Strong understanding of reinforcement learning: policy optimization, reward design, exploration, and environment design
Experience with model adaptation including fine-tuning workflows, dataset construction, and running models under tight compute/memory budgets
Proficiency with quantization techniques (PTQ and QAT), pruning, and architecture search for efficient inference
Hands-on experience building evaluation infrastructure and automated scoring systems
High agency, ownership-driven mindset—ship without being asked, solve hard problems because you want the answer

Work arrangement
No

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App