Live opening · Posted 7 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Job Description: ML Data and Evaluation Engineer / A2
Position Title: ML Data and Evaluation Engineer
Location: Remote
Job Type: Full Time
Experience: Senior position 8+ years
Junior position 4+ years
Job Summary: We are seeking ML Data and Evaluation Engineer to turn raw test execution logs into reliable model training data and build the measurement pipeline used to decide whether model changes are safe and useful.
Job Responsibilities:
● Parse logs into structured supervised examples containing context, candidate tools, expected actions and reference labels where available.
● Deduplicate, balance and version datasets; cover difficult and long-context tasks; implement privacy, schema and label quality checks.
● Create training and validation splits that avoid task or session leakage; support an agreed independent final evaluation set.
Qualifications:
● Build the evaluation harness, error slices and disagreement reports; track dataset and model versions so results are reproducible.
● Strong Python and SQL, robust data pipelines and automated validation; hands-on ML or LLM dataset preparation.
● Applied deduplication, sampling, annotation quality and dataset splitting; can identify leakage and biased coverage in repeated logs.
● Experience building model or task evaluation code, testing metric correctness and tracing unexpected results back to inputs and labels.
● Useful additional experience: AWS S3, Spark for larger datasets, experiment tracking, tool-calling schemas and browser test execution data. Familiarity with model training is useful; this profile primarily owns data and measurement, direct ownership of a completed model experiment. A versioned dataset with measurable quality gates and a reproducible benchmark report, including difficult and long-context examples. Failures should be traceable to a task, dataset version and model version.
What to include with your CV Describe a training dataset or evaluation pipeline you built, its size and quality issues, your split and deduplication approach, one misleading metric you corrected and the resulting improvement. Include role code, current Indian city, exact earliest start date, interview availability and expected engagement terms. Delivery is alongside AWS Professional Services and enterprise engineering teams; customer details are confidential.
What We Offer:
● Competitive salary and benefits package.
● Leadership opportunities and a path for career advancement.
● A collaborative and innovative work environment.
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.