Live opening · Posted 2 days ago

AI Engineer

BusinessOnBot · Bangalore
Instahyre 3-5 yrs
You are 2 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 2 days ago
CompanyBusinessOnBot
LocationBangalore
Experience3-5 yrs
SourceInstahyre
Listed2 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
14 min from Instahyre publishing this role to us finding it
3 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
16,165 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

We're looking for an AI Engineer who can own our agent stack end-to-end: design, build, ship, and keep it healthy in production. This is not a research role. You won't be reading papers and running benchmarks in a notebook. Our AI agents talk to real customers of real D2C brands, on WhatsApp and on voice, about their orders and their money. When an agent hallucinates a discount code or loops on a tool call, a business owner notices and a customer is unhappy. So the job is to build agents that work reliably, prove they work with evals and traces, keep them from doing dumb or dangerous things with guardrails, and do all of it without the inference bill quietly doubling every quarter. You'll be the person who holds the entire agent architecture in their head. That means being independent; you'll get direction on what the product needs, not a spec for how to build it.
The core responsibilities for the job include the following:
Agent Architecture, End-to-End:
Own the design and implementation of the production LLM agent tool calling, memory, state, multi-step reasoning, and handoff to humans.
Build and maintain Python/FastAPI services that expose these agents to our platform, reliably and at low latency.
Decide agent boundaries: what's one agent vs. many, what's a tool vs. a workflow step, and what should never be an LLM call at all.
Integrate with our core backend, messaging channels, and e-commerce/CRM data sources so agents actually have the context they need.
RAG and Retrieval Quality:
Build and own retrieval pipelines for chunking, embedding, indexing, reranking, and refresh strategy.
Debug retrieval failures properly: is it the chunking, the query, the index, or the prompt? Know how to tell the difference.
Handle multi-tenant data isolation in retrieval as a correctness and security requirement, not an afterthought.
Keep retrieval fresh as brand catalogs, policies, and offers change.
Guardrails and Safety:
Design input and output guardrails, prompt injection defense, PII handling, scope enforcement, refusal behavior, and escalation paths.
Prevent the failure modes that cost customers money: hallucinated commitments, wrong order info, runaway tool loops, and unbounded retries.
Build deterministic fallbacks for when the model is wrong, slow, or unavailable.
Treat "the agent said something it shouldn't have" as a production incident, not a prompt-tuning ticket.
Observability, Tracing, and Evaluation:
Instrument everything: traces, spans, token usage, latency, and tool call outcomes (we use LangSmith, New Relic, Sentry, and Pino structured logging).
Build eval sets and regression suites so we know whether a prompt or model change made things better or worse, with evidence.
Make agent behavior debuggable by someone who isn't you; a trace should tell the story without a walkthrough.
Set up alerting for quality degradation, not just uptime.
Cost Consciousness:
Own the LLM and infrastructure cost of everything you ship (AWS, OpenAI, Gemini, vector store, embedding compute).
Know the unit economics: cost per conversation, per resolution, and per brand, and be able to quote them.
Make the routing calls smaller model vs. larger, cached vs. fresh, and retrieval vs. long context, and defend them with numbers.
Flag cost anomalies early and propose architectural fixes, not just usage caps.
Customer Empathy:
Understand that our agents are the front door for D2C businesses talking to their customers.
Read real conversations. Sit with the support and product teams. Let actual failures shape the roadmap, not intuition.
Treat a customer-facing quality bug with more urgency than an internal refactor.
Balancing Tradeoffs:
Navigate the tension between shipping a working agent this month and building the platform that supports fifty of them.
Know when a prompt is the right answer and when it's a hack that will break in three weeks.
Make architecture decisions that hold up as the AI surface area grows agents, voice, and deeper e-commerce automation.
Requirements:
3-5 years of software engineering experience, with meaningful time spent shipping LLM-powered features to production.
Strong Python, with production experience in FastAPI (or equivalent async Python service frameworks).
Hands-on experience building agents with LangChain (or LangGraph / similar orchestration frameworks), not just calling a chat completion endpoint.
Practical experience with LangSmith or comparable tooling for tracing, evals, and prompt/version management.
Direct experience with OpenAI and Gemini model families, including their tool calling, structured output, and failure quirks.
Built and operated a RAG pipeline in production, and can explain what broke and how you fixed it.
Real experience designing guardrails and safety layers around LLM output.
Comfort with AWS and with reasoning about cost/performance tradeoffs in cloud and inference spend.
The ability to hold an entire agent architecture independently; you can go from a fuzzy product goal to a shipped, monitored system without hand-holding.
Strong Plus:
Experience building voice AI pipelines: STT/TTS, streaming, turn-taking, barge-in, and latency budgets.
Exposure to real-time messaging systems (WebSockets, pub/sub, queue-driven architectures).
Experience with multi-tenant SaaS architecture and per-tenant data isolation.
Familiarity with WhatsApp Business API, omnichannel messaging, or e-commerce integrations.
Experience working alongside a TypeScript/Node.js backend.
Fine-tuning, distillation, or model routing work done for cost or latency reasons
Prior experience at a startup in the 10-100 person range where you set the standard rather than inherited it.

Experience
3-5 yrs

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App