Live opening · Posted 11 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
About the role
The AI/ML Platform pod builds company' internal AI platform: the gateway, agent runtime, registry, retrieval, and evaluation stack that every team building with LLMs and agents runs on.
As the platform's first dedicated quality analyst, you own the day-to-day evidence of whether our agents actually work. You run the evaluation sets, judge outputs against rubrics, keep the ground truth honest, and turn what you find into concrete fixes for the pods that own each agent.
Responsibilities
• Run evaluation sets against production and pre-release agents on a regular cadence and on every meaningful prompt, model, tool, or retrieval change; report pass rates, regressions, and trends.
• Review agent outputs against defined rubrics: correctness, grounding, tool-use accuracy, safety, tone, and format. Score consistently, document edge cases, and flag rubric gaps.
• Curate and maintain ground-truth datasets: write and verify golden answers, label failures by root cause (prompt, retrieval, tool, model), and retire stale/ambiguous cases.
• Calibrate LLM-as-judge scoring against human review and report where automated judges disagree with people.
• Feed findings back to agent owner pods with reproducible examples, severity, and suspected cause; track fixes through re-evaluation to closure.
• Convert production failures and user corrections into new eval cases so the suite grows with real usage.
Qualifications
• 3+ years in QA, data annotation, analytics, or a similar evidence-driven role; 1+ year hands-on with LLM or chatbot outputs.
• Sharp, consistent judgment when grading open-ended text; comfortable defending a score with a rubric citation.
• Working Python and SQL to run eval scripts, slice results, and pull traces. Familiarity with Langfuse, Arize Phoenix, or similar tracing and eval tools is a plus.
• Understands how RAG, tool calling, and agent loops fail, and can tell a retrieval miss from a model hallucination.
• Writes clear, concise bug reports and summaries for engineers and product owners.
• Fluent written and spoken English.
• Available for regular overlap with US working hours for pod syncs
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.