Live opening · Posted 10 hours ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Help a top AI lab evaluate and improve large language models through security-focused coding tasks. Bring your software-engineering judgment and hands-on security experience to work involving vulnerabilities, exploit verification and security patches.
This is a contracting engagement, with potential for a longer-term engagement. Remote candidates in the selected countries are elegible.
What you will do
Evaluate coding tasks involving software vulnerabilities, exploit verification and security patches.
Create high-quality coding prompts and reference answers for benchmark-style problems.
Evaluate model outputs for code generation, refactoring, debugging and implementation.
Identify and document model failures, edge cases and reasoning gaps.
Compare private language models with leading external models.
Build or configure coding environments for evaluation and reinforcement learning.
Follow detailed annotation and evaluation guidelines consistently.
What you bring
At least five years of professional software-development experience and strong Python skills.
Hands-on experience with vulnerability research, exploit reproduction or verification, or implementing, backporting or validating security patches.
The ability to apply structured evaluation criteria and write clear technical feedback.
Fluency in written and spoken English.
Helpful, not required
Professional code review, coding annotation, LLM/code evaluation or benchmark design.
Knowledge of another programming language.
Team leadership or mentoring experience.
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.