Live opening · Posted 7 days ago

LLM Fine Tuning and Evaluation Engineer

Talentgigs · India (Remote)
Linkedin Yes
You are 7 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 7 days ago
CompanyTalentgigs
LocationIndia (Remote)
Work modeYes
SourceLinkedin
Listed7 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
12 min from Linkedin publishing this role to us finding it
9 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
71,617 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

Job Description: ML Data and Evaluation Engineer / A2
Position Title: ML Data and Evaluation Engineer
Location: Remote
Job Type: Full Time
Experience: Senior position 8+ years
Junior position 4+ years
Job Summary: We are seeking ML Data and Evaluation Engineer to turn raw test execution logs into reliable model training data and build the measurement pipeline used to decide whether model changes are safe and useful.
Job Responsibilities:
● Parse logs into structured supervised examples containing context, candidate tools, expected actions and reference labels where available.
● Deduplicate, balance and version datasets; cover difficult and long-context tasks; implement privacy, schema and label quality checks.
● Create training and validation splits that avoid task or session leakage; support an agreed independent final evaluation set.
Qualifications:
● Build the evaluation harness, error slices and disagreement reports; track dataset and model versions so results are reproducible.
● Strong Python and SQL, robust data pipelines and automated validation; hands-on ML or LLM dataset preparation.
● Applied deduplication, sampling, annotation quality and dataset splitting; can identify leakage and biased coverage in repeated logs.
● Experience building model or task evaluation code, testing metric correctness and tracing unexpected results back to inputs and labels.
● Useful additional experience: AWS S3, Spark for larger datasets, experiment tracking, tool-calling schemas and browser test execution data. Familiarity with model training is useful; this profile primarily owns data and measurement, direct ownership of a completed model experiment. A versioned dataset with measurable quality gates and a reproducible benchmark report, including difficult and long-context examples. Failures should be traceable to a task, dataset version and model version.
What to include with your CV Describe a training dataset or evaluation pipeline you built, its size and quality issues, your split and deduplication approach, one misleading metric you corrected and the resulting improvement. Include role code, current Indian city, exact earliest start date, interview availability and expected engagement terms. Delivery is alongside AWS Professional Services and enterprise engineering teams; customer details are confidential.
What We Offer:
● Competitive salary and benefits package.
● Leadership opportunities and a path for career advancement.
● A collaborative and innovative work environment.

Work arrangement
Yes

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App