Live opening · Posted 1 day ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
ABOUT:
Owns quality for AI-native applications — functional testing plus the AI-specific evaluation (accuracy, drift, hallucination).
KEY RESPONSIBILITIES
Build test plans and automation for AI-native application features (functional + AI-specific)
Design evaluation harnesses for model/agent outputs — accuracy, consistency, hallucination rate
Run regression testing across model/prompt/config changes to catch silent quality drift
Red-team AI features for edge cases and adversarial inputs where relevant
Build automated eval pipelines integrated into CI/CD
Partner with AI Architects to define testability requirements before build starts
Own the quality gate before any AI feature ships to production
Communicate quality risk to delivery leadership in terms they can act on
Train delivery teams on AI-specific testing practices
Own the evals framework for the practice — golden datasets, scoring rubrics, LLM-as-judge calibration, and versioned benchmarks per use case
Define eval acceptance thresholds per engagement and gate releases on them
Build eval engineering tooling — dataset curation, trace capture, offline/online eval runs, and dashboards delivery teams can read
Instrument production evals and drift monitoring, feeding failures back into the golden datasets
REQUIREMENTS & SKILLS
5–9 yrs QA/test engineering, with 2+ yrs testing AI/ML-powered features specifically
Strong test automation skills (Python-based frameworks, CI/CD integration)
Understands AI-specific failure modes — hallucination, bias, drift, non-determinism — and designs tests for them
Statistically literate enough to interpret model evaluation metrics, not just pass/fail results
Familiarity with red-teaming methodologies for AI systems
Clear, assertive communicator — willing to block a release over a quality concern
Detail-oriented and methodical under delivery-timeline pressure
Collaborative but independent — doesn't rubber-stamp under delivery pressure
Explains quality risk in business-impact terms, not just technical jargon
Hands-on evals engineering — builds and maintains eval suites with frameworks such as OpenAI Evals, Ragas, DeepEval, LangSmith, Azure AI Foundry evaluations
Designs golden datasets and rubrics, and calibrates LLM-as-judge scoring against human review
Understands RAG and agent eval metrics — groundedness, retrieval precision/recall, task completion, tool-call correctness, cost/latency
Experience wiring evals and drift monitoring into CI/CD and production observability
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.