Live opening · Posted 1 day ago

AI Platform Engineer

Demandbase · Hyderabad
Instahyre 5-6 yrs
You are 1 day behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 1 day ago
CompanyDemandbase
LocationHyderabad
Experience5-6 yrs
SourceInstahyre
Listed1 day ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
3 min from Instahyre publishing this role to us finding it
5 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
18,181 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

Responsibilities:
The LLM gateway. The single front door to every model provider we use: keys, routing, rate limits, failover, provider quotas. When it's down, every AI feature at Demandbase is down.
Reliability, for real. SLOs, on-call, incident response, and the postmortems for the runtime. This is a genuine SRE ownership role, not "build it and let someone else run it. "
Cost control. Per-team and per-model attribution, budgets, quotas. LLM spend is the kind of line item that quietly triples; your job is to make it legible and then make it smaller.
LLM observability. Traces, spans, prompt/response capture, ingest and billing visibility for every AI feature in the company.
Evals and experimentation infra. The shared harnesses, datasets, and scoring plumbing product teams use to know whether a prompt or model change actually made things better.
Caching and guardrails. Response and semantic caching to cut latency and redundant spend; input/output safety, PII handling, and policy enforcement at the gateway.
Requirements:
Strong production infrastructure background: Kubernetes, AWS, Terraform, GitOps. You've operated systems that people were paged for, and you've been the one paged.
Python, at a level where you're comfortable owning services and tooling in it, not just scripting.
Real on-call and incident-response experience. You can talk about an incident you ran, what you got wrong, and what you changed afterwards.
Observability fluency beyond "we have dashboards": you've instrumented systems, chased cardinality and cost in a metrics/logging backend, and built alerts that fire when they should.
Hands-on familiarity with how LLM systems actually work. You don't need production ML experience. But you need to have built something with these tools an agent, a RAG pipeline, an internal tool and to have used evals or experiments to decide whether it was any good. Tokens, context windows, prompt/response tracing, and why an eval suite is fundamental should all be familiar ground. Candidates who have only read about this are not a fit.
Nice to have:
An LLM gateway or proxy (LiteLLM or similar) in production.
LLM observability tooling: Datadog LLM Observability, LangSmith, Braintrust, Arize, or similar.
Eval frameworks and LLM-as-judge scoring in a real workflow, not a demo.
FinOps instincts: cost attribution, showback, quota design.
Guardrails, PII detection, or content-filtering systems.
GCP alongside AWS.
Tech stack: AWS - GCP EKS, Flux/GitOps, Karpenter Python, LiteLLM, Datadog (LLM observability, APM), Prometheus, Loki, ClickHouse, Terraform, GitLab CI.

Experience
5-6 yrs

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App