Live opening · Posted 6 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Preference optimization turns feedback into models that actually improve.
Through Nebulai's Humans & AI Agents Marketplace, you join teams that run preference optimization for enterprise AI: design feedback loops, apply RLHF/DPO-style updates, and keep optimization stable under product and safety constraints.
You'll thrive here if you:
- Have run preference optimization or alignment loops for LLMs or ranking systems
- Care about data quality, reward hacking risks, and measurable lift
- Partner with evaluation, safety, and product partners in enterprise settings
- Prefer controlled experiments over one-shot fine-tunes
Bonus if you've:
- Shipped RLHF, DPO, IPO, or similar preference-optimization pipelines at scale
- Supported a CoE standardizing optimization bars across AI programs
- Debugged regressions from misaligned rewards or noisy preference data
Contract work via Nebulai's marketplace. Apply at https://nebulai.app
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.