Live opening · Posted 21 hours ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Job Title: Senior Inference Engineer
About the Role
We’re working with a well‑funded AI infrastructure company over 2B in funding building a next‑generation platform for production LLM inference. We’re hiring a customer‑facing inference engineer to help enterprise customers deploy, optimize, and troubleshoot LLM inference workloads in real production environments.
This role is not focused on GenAI application development — it is focused on LLM inference performance, GPU optimization, and runtime expertise.
What You’ll Do
Work directly with enterprise customers on LLM inference deployments
Optimize inference performance (latency, throughput, GPU utilization)
Troubleshoot runtime issues across model, GPU, and infrastructure layers
Support customer adoption of modern inference stacks
Collaborate with product + engineering to improve inference tooling
Must‑Haves
Strong experience with LLM inference in production
Hands‑on expertise with at least one:
vLLM
SGLang
TensorRT‑LLM
Triton
PyTorch Dynamo
GPU inference deployments
Model bring‑up
Solid Python engineering skills
Experience working directly with customers or external engineering teams
Ability to debug complex ML systems end‑to‑end
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.