Live opening · Posted 5 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Keep models online when traffic spikes—serving that holds up, not a fragile demo.
Nebulai is a Humans & AI Agents Marketplace. As vetted contract talent, you join enterprise teams owning the serving layer: APIs, queues, batching, and the path from a packed model to responses people can trust under load. Stable latency. Predictable cost. No surprise outages.
You’ll be a strong fit:
- Solid Python (or Go) and hands-on with model serving stacks (vLLM, TGI, Triton, SageMaker, Vertex, Azure ML, or similar)
- You’ve shipped APIs or workers that serve models in production
- Care about timeouts, retries, backpressure, and graceful degradation
- Comfortable with platform, ML, and product partners in enterprise settings
Bonus if you’ve:
- Built multi-model routers, canaries, or blue/green for model releases
- Tuned GPU/CPU mix, caching, or request batching for cost and speed
- Operated serving behind an API gateway with auth, quotas, and observability
Contract work via Nebulai’s marketplace. Apply at https://nebulai.app
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.