Live opening · Posted 16 hours ago

Senior AI Evaluation Engineer | Remote |

Haparz · India (Remote)
Linkedin Yes
You are 16 hours behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 16 hours ago
CompanyHaparz
LocationIndia (Remote)
Work modeYes
SourceLinkedin
Listed16 hours ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
9 min from Linkedin publishing this role to us finding it
11 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
66,054 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

Senior AI Evaluation Engineer / Staff AI Evaluation Engineer
Experience - 6+ Years
Role Overview
We are looking for a Senior/Staff AI Evaluation Engineer to lead the benchmarking, validation, quality, and safety evaluation of next-generation Agentic AI applications.
This role sits at the intersection of AI/ML, Software Engineering, Quality Engineering, and Test Automation. The ideal candidate should have strong hands-on experience building automated evaluation frameworks and testing LLM, RAG, LangGraph, LangChain, and multi-agent applications.
The candidate will work closely with Engineering, Product, Data, AI/ML, and Quality Engineering teams to establish measurable AI quality standards and build automated regression and evaluation suites for production AI systems.
Key Responsibilities
Design and implement automated evaluation frameworks for LLM, RAG, and Agentic AI applications.
Develop measurable evaluation criteria for AI quality, accuracy, relevance, reliability, safety, and consistency.
Build and maintain golden datasets, benchmark datasets, test datasets, and evaluation suites.
Perform functional, regression, integration, end-to-end, adversarial, and safety testing of AI applications.
Evaluate LangGraph-based and multi-agent applications at the node, state-transition, routing, and workflow levels.
Build automated evaluation workflows using Python and modern testing frameworks.
Use LangSmith for tracing, debugging, dataset management, experiments, and AI evaluations.
Design and execute adversarial testing and identify failure modes, hallucinations, prompt vulnerabilities, and unexpected agent behavior.
Evaluate AI applications for Responsible AI, security, privacy, and safety considerations.
Work with LangChain and LangGraph, including sub-graphs and conditional routing.
Integrate AI evaluation and automated testing into GitHub-based CI/CD workflows.
Work with AWS and Amazon Bedrock for AI application testing and evaluation.
Analyze LLM outputs, embeddings, vector search, prompts, RAG pipelines, and agent behavior.
Develop automated regression suites to detect model, prompt, retrieval, and application-level quality degradation.
Collaborate with engineering teams to troubleshoot failures and improve AI system reliability.
Document evaluation methodologies, test results, quality metrics, and identified risks.
Independently investigate ambiguous technical problems and drive them toward measurable solutions.
Required Qualifications
6+ years of experience in Software Engineering, Quality Engineering, Test Automation, AI Engineering, Machine Learning, or a related technical field.
Hands-on experience testing or evaluating LLM, NLP, Machine Learning, or AI applications.
Strong hands-on Python programming and automation experience.
Experience creating and maintaining golden datasets, benchmark datasets, or test datasets.
Hands-on experience with adversarial testing.
Knowledge of AI Safety, Responsible AI, Security, and Privacy evaluation.
Strong structural understanding of LangChain, LangGraph, and LangSmith.
Hands-on experience with LangSmith for tracing, debugging, experiments, datasets, and evaluations.
Strong GitHub experience including repositories, branching, pull requests, code reviews, and CI/CD workflows.
Experience with Claude Code or similar AI-assisted development tools.
Experience with AWS and Amazon Bedrock.
Strong knowledge of REST APIs, JSON, SQL, and modern application architectures.
Strong understanding of LLMs, prompt engineering, embeddings, vector search, RAG, and AI agents.
Experience with automated, regression, integration, and end-to-end testing.
Hands-on experience evaluating LangGraph-based or multi-agent applications at node and state-transition levels.
Strong analytical, troubleshooting, communication, and problem-solving skills.
Ability to work independently and take ownership of complex and ambiguous technical challenges.
Primary Mandatory Skills:
6+ years of relevant Software/QA/AI Engineering experience
Python
LLM / GenAI Application Testing & Evaluation
LangChain
LangGraph
LangSmithAI/LLM Evaluation Frameworks
Golden / Benchmark Dataset Creation
Adversarial Testing
RAG
Prompt Engineering
AI Agents / Multi-Agent Systems
Automated Testing & Regression Testing
Secondary Mandatory Skills:
AWS
Amazon Bedrock
GitHub & CI/CD
REST APIs
JSON
SQL
AI Safety / Responsible AI
Security & Privacy Evaluation
Embeddings & Vector Search
Integration & End-to-End Testing
Claude Code or similar AI-assisted development tools
Preferred Candidate Profile
Candidates with hands-on experience in AI evaluation engineering, GenAI quality engineering, LLM testing, Agentic AI testing, LangGraph evaluation, LangSmith, RAG evaluation, and automated AI regression frameworks will be particularly relevant for this position.

Work arrangement
Yes

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App