Live opening · Posted 13 hours ago

QA Engineer – Generative AI & Agentic AI

UST · Trivandrum, Kerala, India (On-site)
Linkedin No
JobBeeper subscribers received an alert for this role.

At a glance

The key details from the original listing.

Posted 13 hours ago
CompanyUST
LocationTrivandrum, Kerala, India (On-site)
Work modeNo
SourceLinkedin
ListedPosted 13 hours ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
5 min from Linkedin publishing this role to us finding it
20 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
61,801 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

Role Description
We are looking for an experienced QA Engineer – Generative AI & Agentic AI to lead quality assurance activities for GenAI and Agentic AI solutions.
This role goes beyond traditional functional testing. The successful candidate will be responsible for validating the correctness, reliability, quality, and production-readiness of AI-generated outputs, including multi-step agent workflows, Text-to-SQL pipelines, Retrieval-Augmented Generation (RAG) systems, and LLM-based applications.
The ideal candidate will have hands-on, production-grade experience testing GenAI/Agentic AI applications end-to-end and should have successfully supported at least one production release of a GenAI or Agentic AI solution.
Key Responsibilities
Design, develop, and execute comprehensive test strategies for GenAI and Agentic AI applications, covering functional, non-functional, integration, and AI output-quality testing.
Test multi-step agentic workflows, validating agent orchestration, tool/function calls, task completion, chained decision-making, reasoning consistency, and outputs across individual agent hops.
Evaluate AI-generated outputs for correctness, relevance, coherence, hallucination, consistency, and reliability.
Test Text-to-SQL workflows, including validation of:
Natural language to SQL generation accuracy
SQL syntax and logic
Schema/table/column alignment
Query execution
Result accuracy against Snowflake data models
Validate RAG pipelines, including:
Retrieval relevance
Context grounding
Chunking effectiveness
Context utilization
Faithfulness of generated responses to retrieved information
Apply GenAI evaluation metrics and methodologies such as:
Embedding similarity
ROUGE
BLEU
LLM-as-a-Judge
Other relevant GenAI quality and accuracy metrics
Design and maintain LLM-as-a-Judge evaluation frameworks and scoring rubrics for scalable and automated assessment of GenAI outputs.
Create and maintain golden datasets, evaluation datasets, regression suites, and reusable AI testing frameworks.
Develop test harnesses/scripts for batch evaluation and automated quality validation of LLM-generated outputs.
Own QA activities for end-to-end production releases, including test planning, execution, regression testing, release quality assessment, sign-off, and post-production monitoring.
Partner closely with Data Science, Platform Engineering, Product, and other stakeholders to establish acceptance criteria, quality gates, and go/no-go release benchmarks.
Identify LLM-specific edge cases and failure modes, including hallucinations, prompt injection risks, bias, inconsistent responses, and agent/tool failures.
Document and track defects through resolution and provide clear evidence of AI application quality.
Maintain traceability of test coverage, evaluation results, quality metrics, and release reports for stakeholders and audit requirements.
Mandatory Skills & Experience
GenAI / Agentic AI Testing
Strong, Demonstrable Hands-on Experience With
Generative AI / LLM application testing
Agentic AI and multi-step agent workflows
Agent orchestration and tool/function-call validation
Text-to-SQL testing
Retrieval-Augmented Generation (RAG) testing
LLM-as-a-Judge evaluation
Prompt and LLM behavior validation
Hallucination and response-grounding assessment
Context-window and context-handling concepts
AI output-quality evaluation beyond traditional pass/fail testing
Platforms
Dataiku: Strong hands-on experience is highly preferred and considered the first preference
Exposure to platforms such as OpenAI workspace, Leena AI, or similar GenAI/Agentic AI platforms
Data / Database
Strong hands-on working experience with Snowflake
Ability to understand and validate data models, schemas, generated SQL queries, and query results
AI Evaluation
Practical Experience With At Least Some Of The Following
Embedding Similarity
RO
Provide your feedback on BizChat
Skills
Generative AI, Test Automation, Agentic AI, BLEU, LLM-as-a-Judge, RAG, Snowflake, Text-to-SQL

Work arrangement
No

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App