Live opening · Posted 4 days ago

Data Scientist, Agent Evaluations & Quality

Clera · Palo Alto
Ashby Yes FullTime
You are 4 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 4 days ago
CompanyClera
LocationPalo Alto
Job typeFullTime
Work modeYes
SourceAshby
Listed4 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
16 min from Ashby publishing this role to us finding it
8 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
72,942 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

ABOUT THE ROLE
This role sits at the intersection of applied data science and AI product quality for a small, fast-moving AI productivity startup building autonomous agents that handle email, calendar, browser, and business software tasks. You will own the measurement of agent quality end-to-end: turning ambiguous product behavior into rigorous, actionable evaluation systems that directly guide engineering and product decisions.
WHAT YOU'LL DO
- Architect and maintain automated evaluation pipelines that measure agent quality across capabilities and product surfaces.
- Translate agent capabilities into explicit success criteria, including pass, partial-pass, and failure definitions for complex multi-step tasks.
- Build representative gold datasets and regression suites covering common workflows, edge cases, ambiguous requests, and adversarial scenarios.
- Define and track metrics such as task success, tool-selection accuracy, instruction adherence, factual consistency, latency, cost, and reliability.
- Design deterministic and model-based graders, calibrate LLM-as-a-judge systems, and measure grader agreement, false positives, and false negatives.
- Analyze traces, tool calls, model outputs, and production outcomes to identify root causes and build a useful failure taxonomy.
- Compare models, prompts, tools, and capability implementations using rigorous offline experiments and production evidence.
- Build dashboards and release-quality signals that make evaluation results understandable and actionable for engineering, product, and leadership.
- Partner with capability engineers to recommend improvements and verify that fixes raise quality without unacceptable regressions in cost, latency, or reliability.
WHAT WE'RE LOOKING FOR
- 5+ years in data science, machine learning, or analytics roles, with a focus on evaluation systems, metrics frameworks, or quality measurement for production systems.
- Demonstrated experience designing and implementing evaluation frameworks, grading systems, and success criteria for ML or AI systems in production.
- Strong Python and SQL proficiency with the ability to build automated data pipelines and production-quality analysis code at scale.
- Solid statistical and experimental design knowledge: sampling, variance, uncertainty quantification, bias detection, confounding variables, and significance testing for non-deterministic systems.
- Experience with ground-truth data development: labeling guideline design, annotation quality control, ambiguity resolution, and dataset maintenance.
- Working knowledge of LLM behavior, tool use, retrieval systems, multi-step execution, and practical failure modes of language model systems.
- Ability to connect quantitative patterns to individual system traces and identify failure origins across model, prompt, context, tools, data, and application logic.
- Experience communicating evaluation results, methodology, uncertainty, and trade-offs to both technical and non-technical stakeholders.
- Comfort operating with high ownership in ambiguous, fast-moving environments, independently turning open-ended quality questions into evaluation systems.
- Experience with LLM-as-a-judge systems, agentic or multi-step task evaluation, or benchmarking platforms for AI systems is a strong plus.
LOCATION
On-site in Palo Alto, California, United States. Visa sponsorship is not available for this role.

Employment type
FullTime

Work arrangement
Yes

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App