Live opening · Posted 5 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Company Description
Founded in 2024 in London, BeatpulseLabs is building the perception layer of AI models and products. We create high-fidelity, custom multimedia AI training datasets for enterprise clients and frontier labs, focusing on safe and domain-specific AI adoption.
Our platform transforms human intelligence, judgement and taste into actionable signals for machine learning systems, using purpose-built annotation schemas and continuous feedback loops.
In the last nine months, we have scaled to more than 5,500 dedicated subject matter experts, and we work with three of the MAG7, as well as most of the leading frontier labs.
We are now building an RL environments practice. We already have the expert network. We are hiring the person to build the technical foundation on top of it.
The role
This is the lead technical role in our RL environments practice. You will own the architecture, the technical direction and the quality bar for every environment we ship. You will work directly with research teams at frontier labs, turn their capability goals into environment and reward designs, and build the team that delivers them. You will lead and mentor that team of engineers and research engineers, and be responsible for its culture and standards. The mandate is broad: you will define how BeatpulseLabs builds RL environments, and your decisions will shape a core business line.
Requirements
Hands-on experience designing RL environments, reward systems or evaluation harnesses used to post-train or evaluate large models, including diagnosing reward hacking, noisy signals, and tasks that are unsolvable or trivial
Practical depth in RLHF, RLVR and reward modelling, including how reward design choices show up in trained model behaviour
Experience turning human or expert judgement into scalable signals: rubrics, grader calibration, LLM judges and inter-rater agreement
8+ years in software or ML engineering, including owning production systems from design through operation at scale
Experience building large-scale data or execution pipelines in the cloud, including containerised and sandboxed workloads
A track record of hiring and leading engineers, and setting technical direction for a team
Experience working directly with external research or technical customers, from scoping through delivery
Why this role
Founding ownership of a business line in one of the most consequential areas of AI
A direct line to the research teams at the world's leading AI labs
Access to an expert network of a calibre few companies can match
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.