Live opening · Posted 7 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Reward signals decide whether optimization actually improves the model.
Through Nebulai's Humans & AI Agents Marketplace, you join teams that design and harden reward signals for enterprise AI: define feedback features, detect reward hacking, and keep training objectives honest under product and safety constraints.
You'll thrive here if you:
- Have built reward models or preference/feedback signals for LLMs or ranking systems
- Care about signal quality, gaming risks, and measurable lift
- Partner with evaluation, safety, and product partners in enterprise settings
- Prefer controlled experiments over one-shot fine-tunes
Bonus if you've:
- Shipped RLHF, DPO, or reward-model pipelines at scale
- Supported a CoE standardizing reward/feedback bars across AI programs
- Debugged regressions from noisy labels or mis-specified rewards
Contract work via Nebulai's marketplace. Apply at https://nebulai.app
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.