Live opening · Posted 5 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
PFB JD for AI Test Engineers
Key Responsibilities
AI Response Evaluation & Validation
Evaluate AI/LLM responses using multiple approaches including:
Human Evaluation
Automated Evaluation
LLM‑as‑a‑Judge techniques
Validate AI outputs for accuracy, relevance, grounding, consistency, and completeness.
Design and execute test scenarios for RAG‑based applications, ensuring response faithfulness and reduced hallucinations.
AI Evaluation Frameworks & Tooling
Apply or integrate AI evaluation frameworks such as:
RAGAS
LangSmith (good to have)
Define evaluation metrics (precision, recall, faithfulness, toxicity, bias, etc.) and analyze results.
Collaborate with engineering teams to embed AI evaluation into CI/CD pipelines where applicable.
Responsible AI & Governance
Ensure compliance with Responsible AI guidelines, including:
Fairness
Bias detection
Safety
Transparency
Validate system instructions, prompt templates, and guardrails.
Support system instruction extraction, prompt testing, and jailbreak prevention strategies.
Test Data & Dataset Management
Understand business and functional requirements to:
Design and curate high‑quality datasets (golden datasets, adversarial datasets, edge‑case datasets).
Prepare datasets for training validation, evaluation, and regression testing.
Maintain dataset versioning and traceability.
Collaboration & Quality Advocacy
Work closely with AI engineers, product owners, and domain SMEs to understand AI behavior and risks.
Provide actionable insights on AI quality issues, improvements, and risks.
Contribute to AI testing best practices, standards, and reusable assets.
Required Skills & Qualifications
Mandatory Skills
Strong understanding of AI/LLM testing concepts and AI quality assurance.
Experience in AI response evaluation techniques (human, automated, LLM‑as‑judge).
Ability to analyze AI outputs and identify:
Hallucinations
Bias
Inconsistencies
Hands‑on experience in test data / dataset creation based on requirements.
Awareness of Responsible AI principles and governance models.
Good to Have
Experience with RAG evaluation frameworks such as RAGA
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.