Live opening · Posted 3 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Treat LLM quality as an engineering problem—not a vibe check.
Through Nebulai, you embed with client teams building eval harnesses for RAG, agents, and copilots: golden sets, rubric-based judges, regression suites, online monitoring, and dashboards that tell product when quality slipped. Fewer silent failures. Clearer ship/no-ship calls. Governed delivery that escapes the pilot trap.
You’ll be a strong fit if you:
- Have built offline and/or online evals for LLM apps (faithfulness, task success, safety)
- Are solid in Python and comfortable with experiment tracking and CI hooks
- Can partner with PMs and domain experts to define what “good” actually means
- Know common failure modes: hallucination, retrieval miss, tool misuse, prompt drift
Bonus if you’ve:
- Used frameworks like RAGAS, DeepEval, Braintrust, LangSmith, or custom judges
- Instrumented production tracing for LLM calls
- Worked in regulated settings where auditability matters
Apply at https://nebulai.app
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.