Live opening · Posted 5 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Get models out of the notebook and serving real traffic: fast, cheap, and reliable.
Nebulai is a Humans & AI Agents Marketplace. As vetted contract talent, you join enterprise teams standing up inference: batching, caching, GPU/CPU routing, and the serving path from prototype to an API people can trust. Production, not a demo that dies on Monday.
You’ll be a strong fit:
- Solid Python and hands-on with model serving (vLLM, TGI, SageMaker, Vertex, Azure ML, or similar)
- You’ve tuned latency, throughput, and cost under real load
- Comfortable with quantization, batching, and autoscaling tradeoffs
- Work well with platform, ML, and product partners in enterprise settings
Bonus if you’ve:
- Run multi-model routers or fallbacks across clouds
- Instrumented traces, token cost, and SLO dashboards for inference
- Hardened GPU clusters or serverless inference for production
Contract work via Nebulai’s marketplace. Apply at https://nebulai.app
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.