Live opening · Posted 6 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Feedback optimization turns real user signals into models that keep getting better.
Through Nebulai's Humans & AI Agents Marketplace, you join teams that run feedback optimization for enterprise AI: design feedback capture loops, apply RLHF/DPO-style updates, and keep optimization honest under product and safety constraints.
You'll thrive here if you:
- Have run feedback or preference optimization loops for LLMs or ranking systems
- Care about signal quality, reward hacking risks, and measurable lift
- Partner with evaluation, safety, and product partners in enterprise settings
- Prefer controlled experiments over one-shot fine-tunes
Bonus if you've:
- Shipped RLHF, DPO, IPO, or similar feedback-optimization pipelines at scale
- Supported a CoE standardizing feedback bars across AI programs
- Debugged regressions from noisy feedback or misaligned rewards
Contract work via Nebulai's marketplace. Apply at https://nebulai.app
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.