Live opening · Posted 2 days ago

Research Engineer, Benchmarks

Clera · Singapore
Ashby No FullTime
You are 2 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 2 days ago
CompanyClera
LocationSingapore
Job typeFullTime
Work modeNo
SourceAshby
Listed2 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
9 min from Ashby publishing this role to us finding it
19 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
60,807 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

ABOUT THE ROLE
This is a hands-on research engineering role focused on designing and owning high-quality benchmarks that evaluate frontier AI agents on realistic, domain-specific workflows. You will sit within a small, highly technical team and play a critical part in ensuring evaluations are rigorous, credible, and trusted by leading AI labs and customers.
WHAT YOU'LL DO
- Design, implement, and own the quality of internal benchmarks for evaluating frontier agents on domain-specific tasks.
- Partner with subject-matter experts to define realistic workflows and translate them into benchmark tasks and evaluation criteria.
- Build and operate reliable infrastructure to run models and agents against benchmark tasks at scale.
- Develop metrics and statistical analyses that measure benchmark difficulty, reliability, and failure modes.
- Validate that benchmark performance correlates with real-world evaluations, customer needs, and frontier lab expectations.
- Write clear technical documentation and benchmark reports for research and engineering audiences.
WHAT WE'RE LOOKING FOR
- 2 to 4 years of experience in software engineering, ML engineering, or research roles, with a focused track record in AI benchmarks or evaluation infrastructure.
- Strong proficiency in Python, Docker, and Linux environments.
- Demonstrated experience designing, implementing, and running benchmarks or evaluation environments for AI agents or large language models.
- Experience building infrastructure to reliably run AI models or agents against benchmark or evaluation tasks.
- Ability to analyze and model workflows across diverse technical or business domains to support task design.
- Sharp attention to detail with a habit of spotting subtle inconsistencies and edge cases.
- Comfort reasoning from first principles about task design, scoring, and failure modes.
- Strong written communication skills; experience producing technical documentation or benchmark reports.
- Ability to thrive in unstructured problem spaces at an early-stage startup.
- Bonus: experience with reinforcement learning pipelines, data generation, or RL agent evaluation; published work on AI benchmarking or model evaluation.
COMPENSATION & BENEFITS
Salary range: USD 150,000 to 250,000 annually. Visa sponsorship is available.
LOCATION
On-site in Singapore.

Employment type
FullTime

Work arrangement
No

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App