Live opening · Posted 18 days ago

Data Scientist

Sarvam · Bangalore
Instahyre 3-6 yrs
You are 18 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 18 days ago
CompanySarvam
LocationBangalore
Experience3-6 yrs
SourceInstahyre
Listed18 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
0 min from Instahyre publishing this role to us finding it
6 hours subscribers had this role before this page existed
16,373 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

Responsibilities:
Design and build evaluation frameworks for Sarvam's AI outputs across domain-specific requirements: document comprehension, command summarization, geospatial reasoning, enterprise workflow automation, and others as they emerge.
Define quality metrics in collaboration with domain experts and clients; translate operational requirements into measurable, defensible signals.
Run structured evaluation cycles pre- and post-deployment; build dashboards that surface model quality in production.
Identify failure modes, edge cases, and distribution shifts with the bias of someone looking for what's wrong, not confirming what's right.
Collaborate with the MLOps Engineer to operationalize evaluation pipelines automated, triggered by deployment events, versioned, and reproducible.
Build and manage domain-specific datasets for fine-tuning, evaluation, and benchmarking, including human annotation workflows where needed.
Publish internal findings and quality reports that feed the product and engineering roadmap.
Requirements:
3-6 years in data science, ML research, or applied AI; at least 2 years working with LLMs in production contexts.
Strong statistics and probability fundamentals you understand what makes an evaluation valid and what makes it misleading.
Experience designing evaluation frameworks from scratch: custom metrics, inter-rater reliability, and red-teaming methodologies.
Python proficiency; comfort with pandas, NumPy, Hugging Face datasets, RAGAS, EleutherAI Eval Harness, LangSmith, or equivalent.
Experience with prompt engineering, model fine-tuning, or RLHF in applied settings.
Ability to work with unstructured domain data: PDFs, doctrine documents, transcripts, and field reports.

Experience
3-6 yrs

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App