Live opening · Posted 13 hours ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
We build the evaluation layer that validates AI agents before they reach customers — an automated system that scores across large volumes of agent traces.
You'll own core parts of that platform: the pipeline that runs traces through model-based judges at scale, and the scoring logic that turns raw output into results teams can act on.
What you'll do
Design evaluation methodologies and benchmarks for agent reasoning, planning, tool use, reliability, and safety — across LLM-as-a-Judge, trajectory-based, and human evaluation
Take problems from research question to prototype to shipped feature, owning them end to end
Build and harden the pipelines and scoring logic behind customer-facing evaluation
Curate synthetic and real-world datasets; measure the evaluator itself for consistency and agreement with human labels
What we're looking for
5+ years in ML, applied AI, prompt engineering, agentic AI, including shipping something real to users
Strong Python skills
Practical depth in agentic AI and context engineering: planning, reasoning, memory, tool use, retrieval, long-context
Experience designing evaluation methodologies, not just running evaluations
Hands-on production work with LLM APIs — prompt engineering, structured output, cost and latency tradeoffs
Clear communication with technical and non-technical audiences
Good to have Experience with AI-assisted development tools (Claude code, Windsurf, or similar)
Employment type
Full-time
Work arrangement
Hybrid
More openings worth a look
Recently tracked roles with full details and direct application links.