Live opening · Posted 7 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
This role sits at the intersection of multimodal AI, real-time voice, vision-language models, mobile-use agents, memory, on-device intelligence, and production-grade evaluation systems. You will help build an AI system that can understand the user’s physical context through camera, microphones, device state and memory - then turn that context into useful action.
Minimum 3+ years of professional experience.
Agentic systems - multi-step task execution, planning and decomposition, state management
Apply quantization (PTQ and QAT), pruning, and architecture search to hit per-product size, latency, and power budgets and efficient in working with SLM’s
Applied vision-language models - prompt architecture, structured output, context management, and visual understanding of interfaces, design and execute distillation strategies.
Scale simulation and scaffolding environments for agentic RL: code execution sandboxes, computer use environments, tool-calling harnesses, and Mobile Use Agents.
Evaluation infrastructure - task suites, automated scoring, regression harnesses, and success / latency / cost measurement
Data pipelines - capture, schema design, labelling, privacy-safe handling, and dataset curation for training
Real-time voice - ASR and TTS integration, streaming interaction, and end-to-end latency engineering
Deep understanding of reinforcement learning: policy optimization, reward design, exploration, and the interplay between environment design and agent behavior.
Model adaptation - fine-tuning workflows, dataset construction, and running models under tight compute and memory budgets
Partner with our mobile and hardware engineers to move capability from the cloud onto the device
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.