Live opening · Posted 7 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Reward modeling turns preference signals into a score a model can chase.
Through Nebulai's Humans & AI Agents Marketplace, you join teams that build reward models for enterprise AI: design preference datasets, train scorers that match human judgment, and keep optimization honest under product and safety constraints.
You'll thrive here if you:
- Have trained reward or preference models for LLMs or ranking systems
- Care about label quality, calibration, and reward hacking risks
- Partner with evaluation, safety, and product partners in enterprise settings
- Prefer measurable reward proxies over vibes-based tuning
Bonus if you've:
- Shipped RLHF, DPO, or reward-model pipelines at scale
- Supported a CoE standardizing reward bars across AI programs
- Debugged reward hacking, sycophancy, or misaligned optimization
Contract work via Nebulai's marketplace. Apply at https://nebulai.app
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.