Live opening · Posted 1 day ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Conduct operability gate reviews: assess evidence packs, reproduce evaluation results, and recommend go/no-go with documented findings.
Own the model-update regression pack methodology and adjudicate regression runs against archived baselines with AIML Operations team.
Calibrate gate thresholds against production reality and maintain evaluation drift hygiene (golden-set rotation, hold-out sets, judge calibration).
Produce the monthly quality report per agent: eval trends, failure-mode taxonomy, and defect clusters with reproduction traces.
Mentor members of the team when needed and review their work for quality and consistency.
Bachelor’s or Master’s degree in Computer Science or a related field
6+ years in ML/data/software with strong evaluation or quality focus
Hands-on experience evaluating LLM or ML systems
LLM/agent evaluation design and statistical rigour
Python and evaluation tooling (promptfoo, DeepEval, or custom harnesses)
Data analysis and metric interpretation
Tracing/observability tooling
Analytical rigour and attention to detail
Clear technical writing for gate findings
Ability to influence build teams on quality
Understanding of telco customer intents and journeys
Excellent communication, stakeholder management
Job Stability - min 2 years in an organization
Notice Period - Immediate to 45days
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.