Live opening · Posted 1 day ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
AI Engineer — Agentic AI, RAG & Low-Latency Voice Systems
Mobiloitte Technologies (I) Pvt. Ltd. | New Delhi (on-site) | Full-time | 5+ years
Compensation: No ceiling for the right candidate.
Two things decide this role: how well you handle latency, and how far past basic RAG you can build.
Mobiloitte builds production AI for enterprise clients across India, the US, UK, UAE, Singapore and South Africa. We're hiring a senior AI Engineer to own our agentic AI, retrieval and voice stack — systems that go live, carry real call volume, and stay up.
1. LATENCY IS THE JOB, NOT A DETAIL
Our voice bots handle inbound and outbound calls where a 400ms difference decides whether the conversation feels human. We need someone who treats latency as an engineering discipline:
- Owns end-to-end voice latency from user speech to first token of bot audio — and can quote the numbers they achieved
- Profiles and attacks each hop: STT, retrieval, inference, TTS, network
- Streaming responses, speculative/partial generation, response chunking to shorten time-to-first-audio
- Barge-in and interruption handling that actually works mid-sentence
- Model caching, warm pools, batching, concurrency control under real call load
- Knows when the fix is architectural, not a bigger GPU
If you can't tell us what your p95 latency was and what you did to bring it down, this isn't the role.
2. AGENTIC AI — BEYOND BASIC RAG
Single-shot retrieve-and-answer is table stakes. We're building systems that plan, act and self-correct:
- Agent orchestration — LangGraph, CrewAI, AutoGen, OpenAI Agents SDK, or your own framework (we care that you've built it, not which one)
- Tool and function calling, API/DB actions, structured outputs
- Multi-step reasoning, task decomposition, planning loops with retry and fallback
- Multi-agent patterns — routing, supervisor/worker, handoffs
- Conversation and long-term memory management across sessions
- Advanced retrieval: agentic and query-rewriting RAG, hybrid search, reranking, graph/multi-hop retrieval, chunking strategy chosen deliberately
- Evaluation and guardrails — hallucination reduction, groundedness scoring, tracing and observability (LangSmith, Langfuse or similar)
- Cost control: token budgeting, context-window discipline, model routing
3. VOICE — INBOUND AND OUTBOUND
- Full pipeline: STT → LLM/agent/RAG → TTS
- Whisper / Faster-Whisper / Deepgram; Piper / Coqui / XTTS or equivalent
- SIP and telephony integration, WebSockets, real-time streaming
- Built the stack rather than wrapping Vapi, Retell or Bland end to end
4. PRIVATE / SELF-HOSTED DEPLOYMENT
Many of our clients cannot send data to public APIs. You should be able to:
- Deploy Llama, Qwen or Mistral inside a client VPC or on-prem
- Run vLLM, Ollama or TensorRT-LLM in production
- Size GPUs and VRAM, apply INT8/INT4 quantization, tune inference
- Handle concurrency, autoscaling, secure API exposure, auth and data isolation
- Explain how you'd replace an OpenAI dependency with a self-hosted architecture — this earns strong preference
ON COMPENSATION
We've deliberately not published a band. For an engineer who has genuinely built agentic systems, driven latency down in production, and run models on their own infrastructure, compensation is not the constraint — we'll match what the work is worth.
Be honest with yourself before applying: if your AI work is mainly calling third-party APIs and writing prompts, we'll find that out in the first twenty minutes. If you've built agents, fought latency, and put a model on your own GPU, we want to talk.
HOW TO APPLY
Apply here and include one thing: a link, repo, demo video or short architecture note for a system you personally built — and tell us what you did on it.
careers@mobiloitte.com
#AIEngineering #AgenticAI #RAG #VoiceAI #LLM #Hiring #Delhi
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.