Live opening · Posted 15 days ago

QA Lead- AI Evaluation & Quality

National e-Governance Division · New Delhi, Delhi, India (On-site)
Linkedin No
You are 15 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 15 days ago
CompanyNational e-Governance Division
LocationNew Delhi, Delhi, India (On-site)
Work modeNo
SourceLinkedin
Listed15 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
3 min from Linkedin publishing this role to us finding it
15 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
71,833 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

QA Lead- AI Evaluation & Quality
Educational Qualification
B.Tech./B.E. or M.Tech./M.S./M.Sc. in Computer Science, Data Science, AI/ML, or a related quantitative discipline (Must have)
Advanced degree (M.Tech./M.S./Ph.D.) or published/applied work in ML evaluation, NLP/LLM evaluation, or Responsible AI is highly desirable
Certifications (Desirable): ISTQB Foundation and ISTQB Certified Tester AI Testing (CT-AI); DeepLearning.AI or equivalent ML foundations; a cloud (AWS/Azure/GCP) or MLOps certification
Non-traditional backgrounds with demonstrable model-evaluation or ML-quality work, open-source evaluation or data-quality contributions, or competitive benchmark/Kaggle achievements will be considered in lieu of formal qualification
Experience
5+ years in ML or software quality assurance, test automation, data or ML engineering, or applied AI; minimum 3 years specifically in AI/ML model evaluation, LLM/Gen-AI evaluation, or ML quality assurance
Demonstrable hands-on experience building evaluation harnesses, benchmark/golden datasets and automated evaluation suites for ML or LLM systems
Demonstrable experience evaluating LLM/RAG systems (hallucination, groundedness, retrieval quality) and/or agentic systems (tool-use correctness, trajectory, task completion)
Prior experience embedding evaluations into CI/CD pipelines and versioning datasets, prompts and experiments
Prior experience delivering AI quality or evaluation work in government, BFSI, healthcare, or another regulated sector is a strong plus
Prior experience directing, mentoring, or setting QA/evaluation standards across multiple teams is desirable, given the programme-lead nature of the role
Key Responsibilities
Evaluation Strategy, Standards & Quality Gates
Own the programme’s AI evaluation methodology and make evaluation a formal gate in the delivery lifecycle — defining, publishing and enforcing common metrics, acceptance thresholds and quality gates
Select the appropriate evaluation approach per use case — offline benchmarks, human evaluation, LLM-as-judge, or online/production evaluation — and document the rationale
Ensure Responsible AI evaluation is built into the standard delivery flow rather than added afterwards, aligned to the MeitY Responsible AI advisory and the IndiaAI Safe & Trusted AI framework
Evaluation Infrastructure & Reusable Harnesses
Design and build automated evaluation harnesses, benchmark and golden datasets, and regression suites that pods reuse across ministry deployments — mirroring the programme’s build-once, reuse-many model
Wire evaluations into CI/CD as release gates; version datasets, prompts and evaluation configurations so results are reproducible and audit-ready
Publish reusable evaluation sets, rubrics, harnesses and templates to AIKosh under standard metadata (and to OpenForge where code is shared) for national reuse
Model & Output Quality Evaluation
Evaluate models with task-appropriate quantitative metrics — precision/recall/F1, ROC-AUC/PR-AUC, calibration, MAE/RMSE and task-specific measures — reported with statistical rigour, including uncertainty, adequate sample sizes and vigilance against metric gaming
Evaluate AI outputs against defined quality dimensions — accuracy, groundedness, reliability, consistency, relevance and adherence to business and policy requirements — rather than pass/fail alone
Gen-AI, LLM & RAG Evaluation
Assess LLM and RAG systems for hallucination, factual consistency, groundedness and citation faithfulness, relevance, prompt robustness, input validation and non-deterministic behaviour
Measure retrieval quality and generation quality separately in RAG pipelines using RAGAS, DeepEval, TruLens or LangSmith
Use LLM-as-judge where appropriate, calibrated against human labels, controlling for known judge biases (position, verbosity, self-preference), and defaulting to human evaluation where automated judgement is unreliable
Agentic AI Evaluation
Evaluate agents on their trajectories, not only final answers — tool selection and tool-call correctness, multi-step reasoning and workflow completion, agent handoffs, failure recovery, memory behaviour and end-to-end task-completion accuracy
Validate human-in-the-loop controls — review, escalation, override, feedback and exception handling — wherever AI decisions require human supervision
Safety, Adversarial & Responsible AI Testing
Under the technical guidance of the AI Safety Researcher, operationalise adversarial and red-team findings — jailbreaks, prompt injection, indirect prompt injection, unsafe behaviour and sensitive-data (PII) leakage — into standard, reusable test suites the pods apply, and verify guardrail effectiveness (Llama Guard, NeMo Guardrails, Guardrails AI or equivalent)
Evaluate fairness by measuring performance across demographic, linguistic, geographic and other relevant cohorts using appropriate fairness metrics (Fairlearn, AI Fairness 360)
Author the QA-side safety and quality evidence — model cards, dataset sheets, evaluation logs and Responsible AI records — for approval
Data & Dataset Quality Validation
Validate training, evaluation and inference data for missing values, duplicates, schema conformance, class imbalance, data leakage, distribution/covariate shift, feature consistency and train–serve skew, using tools such as Great Expectations, Evidently or Deepchecks
Curate high-quality evaluation datasets that reflect real citizen usage, edge cases and Indian-language and regional diversity
Production Monitoring & Quality Operations
Monitor live AI applications for quality and retrieval degradation, drift (data, concept and model), accuracy and latency regressions, token/cost behaviour, failures and safety/compliance violations
Define alerts, thresholds and dashboards for model-quality, latency, error-rate, drift and compliance signals; support quality-incident response and feed production findings back into the evaluation sets
Human Evaluation Operations
Design evaluation rubrics and run structured human evaluation and annotation; manage annotator quality and inter-annotator agreement; build preference/label datasets that ground and validate the automated metrics
Cross-Pod Leadership, Mentoring & Governance
Provide technical direction to the AI QA Engineers in the team; set shared standards and harnesses, keep pod-level practice consistent and current, and drive adoption of sound, current evaluation techniques while screening out hype
Turn quality evidence into decision-ready go/no-go recommendations for deployments, and represent AI quality in programme governance and monthly reviews
Contribute AI quality and evaluation requirements to NeGD RFQs and RDRs, and review agency deliverables for evaluation and quality posture
Technical Competencies
Programming & Data: Python (advanced) and SQL; pandas, NumPy, scikit-learn; pytest for test authoring
Evaluation & Measurement: offline benchmarks, human evaluation, LLM-as-judge and online/production evaluation; statistical rigour — significance testing, bootstrap confidence intervals, calibration (ECE) and inter-annotator agreement (Cohen’s / Fleiss’ kappa)
LLM, RAG & Agentic Evaluation: RAGAS, DeepEval, TruLens, LangSmith or Promptfoo; separate retrieval- and generation-quality measurement; trajectory and tool-use evaluation for agents
ML Metrics: precision/recall/F1, ROC-AUC and PR-AUC, calibration, MAE/RMSE and task-specific metrics
Data & Dataset Quality: Great Expectations, Evidently, Deepchecks or TFDV; detection of leakage, drift and covariate shift, class imbalance, schema violations and train–serve skew
Safety & Responsible AI Testing: guardrail testing (Llama Guard, NeMo Guardrails, Guardrails AI), PII-leakage and unsafe-output testing; fairness metrics (Fairlearn, AI Fairness 360); model cards and dataset sheets — applied under the AI Safety Researcher’s guidance
MLOps & Observability: MLflow or Weights & Biases for experiment and dataset versioning; Prometheus, Grafana, OpenTelemetry; LLM observability (Langfuse, Arize Phoenix or equivalent); CI/CD (Jenkins, GitHub Actions)
Test Automation (supporting): pytest as the primary automation stack; Postman or REST-assured for API testing; Selenium or Cypress where UI validation is relevant
Security (AI-aware): PII/sensitive-data exposure, access controls around AI services, model/API misuse and data leakage, in addition to conventional application-security fundamentals (authentication/authorisation, encryption, secure APIs)
Cloud & Platforms: AWS, Azure, GCP and IndiaAI Compute evaluation and testing environments; containerised evaluation (Docker, Kubernetes)
Government Frameworks & Standards: IndiaAI Safe & Trusted AI framework, MeitY Responsible AI advisory, MeitY security guidelines, CERT-In directions, DPDPA 2023
Communication & Leadership: ability to auth

Work arrangement
No

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App