Live opening · Posted 7 days ago

Technical AI Evaluation Analyst (LATAM)

Gramian Consulting Group · Argentina
Workable No
You are 7 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 7 days ago
CompanyGramian Consulting Group
LocationArgentina
Work modeNo
SourceWorkable
Listed7 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
11 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
73,267 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

About Gramian
Gramian Consultancy is a boutique consultancy specializing in IT professional services and engineering talent solutions. With a strong background in software engineering and leadership, we help companies build high-performing teams by matching them with professionals who truly fit their needs.
About the Role
We are seeking a detail-oriented AI Evaluation & Quality Assurance Specialist to review the quality, correctness, and fairness of tasks designed to evaluate AI agents. You will inspect reference solutions, grading logic, execution traces, and generated deliverables to identify task defects, evaluation errors, and unjustified model failures. This role requires strong technical fluency, independent analytical judgment, and the ability to produce clear, evidence-based feedback.
LOCATION: Remote – Latin America (LATAM)
CONTRACT: Hourly Contractor
COMMITMENT: 40 hours per week
TIME OVERLAP: 8 hours of mandatory PST overlap
DURATION: 10 weeks
START DATE: Immediately
Key Responsibilities
Validate task quality by checking instructions, source materials, reference solutions, and evaluation criteria for consistency and completeness.
Review AI agent execution traces, tool calls, and deliverables to determine whether outcomes are justified.
Audit grading logic to identify brittle checks, incorrect expected answers, unsupported rubric criteria, and unfair penalties for valid alternative solutions.
Investigate discrepancies between model performance, grader results, and expected outcomes.
Distinguish genuine model limitations from task defects, grader errors, and environment or tool failures.
Independently assess automated QC findings rather than accepting them without verification.
Document concise, evidence-backed findings and provide actionable, reproducible feedback.
Flag uncertainty and verify that implemented revisions resolve previously identified issues.
Requirements
Comfortable reading and interpreting Python, SQL, shell scripts, structured data, and execution logs.
Experience reviewing technical workflows, software behavior, data outputs, or evaluation logic.
Strong analytical skills, including the ability to verify calculations and reconcile conflicting evidence.
Demonstrated ability to assess the correctness and completeness of technical deliverables.
Strong written English with experience providing clear, specific, and reproducible feedback.
High attention to detail when identifying inconsistencies, missing information, and evaluation defects.

Work arrangement
No

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App