Live opening · Posted 11 days ago

AI Quality Analyst

Brewcode · Hyderabad, Telangana, India (On-site)
Linkedin No
You are 11 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 11 days ago
CompanyBrewcode
LocationHyderabad, Telangana, India (On-site)
Work modeNo
SourceLinkedin
Listed11 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
6 min from Linkedin publishing this role to us finding it
11 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
73,703 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

About the role
The AI/ML Platform pod builds company' internal AI platform: the gateway, agent runtime, registry, retrieval, and evaluation stack that every team building with LLMs and agents runs on.
As the platform's first dedicated quality analyst, you own the day-to-day evidence of whether our agents actually work. You run the evaluation sets, judge outputs against rubrics, keep the ground truth honest, and turn what you find into concrete fixes for the pods that own each agent.
Responsibilities
• Run evaluation sets against production and pre-release agents on a regular cadence and on every meaningful prompt, model, tool, or retrieval change; report pass rates, regressions, and trends.
• Review agent outputs against defined rubrics: correctness, grounding, tool-use accuracy, safety, tone, and format. Score consistently, document edge cases, and flag rubric gaps.
• Curate and maintain ground-truth datasets: write and verify golden answers, label failures by root cause (prompt, retrieval, tool, model), and retire stale/ambiguous cases.
• Calibrate LLM-as-judge scoring against human review and report where automated judges disagree with people.
• Feed findings back to agent owner pods with reproducible examples, severity, and suspected cause; track fixes through re-evaluation to closure.
• Convert production failures and user corrections into new eval cases so the suite grows with real usage.
Qualifications
• 3+ years in QA, data annotation, analytics, or a similar evidence-driven role; 1+ year hands-on with LLM or chatbot outputs.
• Sharp, consistent judgment when grading open-ended text; comfortable defending a score with a rubric citation.
• Working Python and SQL to run eval scripts, slice results, and pull traces. Familiarity with Langfuse, Arize Phoenix, or similar tracing and eval tools is a plus.
• Understands how RAG, tool calling, and agent loops fail, and can tell a retrieval miss from a model hallucination.
• Writes clear, concise bug reports and summaries for engineers and product owners.
• Fluent written and spoken English.
• Available for regular overlap with US working hours for pod syncs

Work arrangement
No

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App