Live opening · Posted 4 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
MailerMen is hiring a QA Engineer – AI Output Evaluation to review and grade the technical answers produced by AI systems, so that the models behind them can be trained on accurate, well-judged human feedback. This is not an ordinary software-testing position. The day-to-day work is evaluation rather than product release testing, and the project is open only to QA specialists who have already been paid to work on human data for AI training, such as annotation, labelling, RLHF, AI response evaluation, model evaluation or rubric-based grading. RLHF, or reinforcement learning from human feedback, is the training method in which people rate or rank a model's answers and those ratings are then used to teach the model what a good answer looks like. Software QA experience on its own, without hands-on human evaluation of AI output, does not meet this requirement.
The core task is to evaluate and rate AI-generated technical outputs against defined quality criteria and rubrics, with particular attention to answers that look correct on the surface but are in fact inaccurate or incomplete. Alongside that, you will design and assess comprehensive test cases for functional, regression and edge-case scenarios, including negative and boundary tests, and review bug reports and test documentation to confirm that issues are reproducible, that the documentation is complete and that severity has been judged correctly. You will identify, isolate and document defects with precise reproduction steps and record them through structured tracking mechanisms.
Much of the value of the work lies in the written record you leave behind. Feedback and annotations have to be detailed and specific enough that a developer can act on them without coming back for clarification, which is why the project asks for written English at B2 level or above, meaning you can set out a technical argument precisely and be understood first time. You will also work with project teams to refine the evaluation guidelines themselves and to improve testing standards and methods as the project runs. This is a high-volume project, so applying a rubric consistently matters as much as individual judgement. No background in building or training AI systems is expected; your QA domain knowledge, together with the paid human-data experience described above, is what the project needs.
This is a contract engagement, worked remotely from anywhere in India. Compensation is ₹1,500 per hour. The role calls for 7+ years of experience. A LinkedIn profile is required and must be provided with your application. No formal degree is required, because practical, demonstrable testing experience takes precedence, and you will need a reliable internet connection and the readiness to begin promptly. Apply through MailerMen with your updated resume and a summary of your relevant experience.
Responsibilities
Evaluate and rate AI-generated technical outputs against defined quality criteria and rubrics.
Identify answers that appear correct but are in fact inaccurate or incomplete, and record why they fail.
Apply the project rubrics consistently across a high volume of items so that ratings stay comparable.
Design comprehensive test cases covering functional, regression and edge-case scenarios.
Assess test cases for coverage, including negative and boundary conditions.
Review bug reports and test documentation to confirm that issues are reproducible and that the documentation is complete.
Check that defect severity has been assessed correctly and correct it where it has not.
Identify, isolate and document defects with precise reproduction steps.
Record defects and evaluation outcomes through structured tracking mechanisms.
Provide detailed, actionable written feedback and annotations so that developers can resolve issues without needing further clarification.
Work with project teams to refine the evaluation guidelines used on the project.
Contribute to the continuous improvement of testing standards and methodologies.
Requirements
7+ years of proven professional experience in software quality assurance as a QA Engineer, SDET, Test Engineer, QA Analyst or a similar role.
Prior paid human data experience supporting AI training, such as annotation, labelling, RLHF, AI response evaluation, model evaluation or rubric-based grading, is a firm requirement for this project.
Software QA experience on its own, without a human-in-the-loop AI evaluation component, does not satisfy that requirement.
Strong command of QA fundamentals, including test case design, bug tracking and regression testing.
Practical experience with both manual and automated testing strategies.
Hands-on experience with automation frameworks and test management tools such as Selenium, Playwright, Cypress, Appium, Postman, Jira, TestRail, Zephyr or BrowserStack, or comparable tools.
Excellent analytical and problem-solving skills with meticulous attention to detail.
Ability to communicate complex findings clearly in written English at B2 level or above, giving specific and actionable feedback.
A LinkedIn profile is required and must be provided with your application.
No formal degree is required, as practical and demonstrable testing experience takes precedence.
No prior experience in building or training AI systems is required, since your QA domain knowledge is what the project draws on.
A reliable internet connection and the readiness to begin promptly.
Skills
Quality AssuranceSoftware TestingTest Case DesignManual TestingAutomation TestingRegression TestingEdge Case TestingBug TrackingBug ReportingTest Management ToolsSeleniumPlaywrightCypressPostmanJiraTestRailHuman DataData AnnotationRLHFModel Output Evaluation
Benefits
The work is done remotely from anywhere in India.
This is a contract engagement.
Compensation is ₹1,500 per hour.
Related searches
QA Engineer Remote Jobs in IndiaAll QA Engineer jobsJobs in India
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.