Live opening · Posted 5 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Keep models fast, efficient, and cost-aware once they leave the lab.
Nebulai is a Humans & AI Agents Marketplace. As vetted contract talent, you join enterprise teams who care about latency, throughput, and spend - not just accuracy on a benchmark. You’ll profile inference, cut waste, and make production AI feel snappy under real traffic.
You’ll be a strong fit:
- Hands-on with Python and production ML/LLM serving
- Experience profiling latency, GPU/CPU use, and token or request cost
- Comfortable trading off quality vs speed with clear metrics
- Worked with product and platform partners in enterprise settings
Bonus if you’ve:
- Tuned batching, caching, quantization, or speculative decoding
- Used SageMaker, Vertex, Azure ML, or self-hosted inference stacks
- Built dashboards that catch regressions before users do
Contract work via Nebulai’s marketplace. Apply at https://nebulai.app
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.