Live opening · Posted 12 hours ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
About Gnani.ai
Gnani.ai is India's leading enterprise voice AI company, building the language infrastructure that powers intelligent voice agents, speech recognition, and text-to-speech systems at scale across 40+ languages. Our models are deployed across government, BFSI, telecom, and enterprise verticals, processing over 30 million voice AI calls a day. Founded in 2016 and backed by Samsung Ventures and Info Edge Ventures, we are one of four companies selected under the IndiaAI Mission to build foundational AI models for India.
Role Overview
We are hiring an LLMOps Engineer to own how our large language models actually run in production — the serving stack, the optimization work that makes it affordable, and the low-level performance engineering that makes it fast. This is a deep infrastructure and performance role, not a wrapper-and-API role.
You will work on self-hosted foundational models served on large-scale NVIDIA GPU clusters, under real-time latency budgets set by live spoken conversations. Voice inference is unforgiving: time-to-first-token is a product feature, not a number on a dashboard, and every millisecond of tail latency is audible to a caller. Your mandate is to hold latency and quality while driving cost per million tokens down.
The work spans three layers: the serving engine (vLLM, SGLang, NVIDIA Dynamo), the model itself (quantization, distillation, speculative decoding, pruning), and the kernels underneath (CUDA/Triton, attention and MoE kernels, profiling and bottleneck analysis). We expect real depth in at least two of the three and genuine curiosity about the third.
Key Responsibilities :
Inference Serving and Platform
• Own the production LLM serving stack across vLLM, SGLang, and NVIDIA Dynamo — including engine selection per workload, with benchmark evidence to justify it
• Tune the serving path end to end: continuous batching, chunked prefill, paged and radix attention, prefix and KV caching, cache-aware and sticky routing, speculative decode integration
• Design and operate disaggregated prefill/decode and multi-node deployments; configure tensor, pipeline, and expert parallelism for both dense and Mixture-of-Experts models
• Run inference on Kubernetes with autoscaling, rolling model updates, canary releases, and tested rollback paths; enforce staging-to-production promotion gates
• Instrument everything that matters: TTFT, inter-token latency, p50/p95/p99, tokens/sec/GPU, KV cache hit rate, queue depth, GPU utilization and MFU, and cost per million tokens
• Build and maintain internal serving APIs, model registries, and deployment tooling so that model updates are routine rather than events
Model Optimization
• Quantization: FP8, INT8, and INT4 weight and activation quantization plus KV cache quantization — calibration set design, per-layer sensitivity analysis, and accuracy recovery, with measured deltas per language and task
• Speculative decoding: draft-model, EAGLE/Medusa-style, and n-gram approaches; tuned for real acceptance rate and end-to-end latency gain under production traffic, not theoretical speedup
• Distillation: build smaller task-specific students from larger teachers for latency-critical paths, held to explicit eval-parity targets
• Pruning and adapters: structured pruning, LoRA/adapter serving, and multi-adapter batching for per-deployment specialization
• Compilation: torch.compile, CUDA graphs, and TensorRT-LLM engine builds — including diagnosing recompilation and dynamic-shape stalls
Kernel and Low-Level Performance
• Profile with Nsight Systems/Compute and the PyTorch profiler; classify bottlenecks as memory-bound, compute-bound, or launch/synchronization-bound, and act on the classification
• Write and tune custom kernels in CUDA and Triton — attention variants, fused MoE dispatch, sampling, quantized GEMM
• Eliminate host-device synchronization stalls, dynamic-shape recompilation, and unoverlapped collectives in distributed serving
• Work fluently with attention and MoE kernel libraries (FlashAttention, FlashInfer, CUTLASS/cuBLAS, expert-parallel dispatch libraries) and know when to use them versus write your own
• Reason from first principles about arithmetic intensity, memory bandwidth, and roofline limits before reaching for a tool
Evaluation, Reliability, and Cost
• Own the optimization regression gate: no optimized build reaches production without an accuracy and behavior evaluation across languages and task types
• Build load-testing harnesses that replay realistic traffic — concurrency, length distributions, burstiness, multi-turn sessions
• Run capacity planning and cost modeling; set and hit targets for cost per million tokens and per concurrent session
• Carry on-call for inference services, write runbooks, and lead blameless postmortems on latency and availability incidents
• Work closely with the training and post-training teams so that serving constraints inform model architecture decisions early, not after the checkpoint lands
Must Have
• 3 to 6 years in ML infrastructure, model serving, or performance engineering, with at least 2 years specifically on LLM inference in production
• Hands-on production experience with at least two of vLLM, SGLang, TensorRT-LLM, or NVIDIA Dynamo, and the ability to explain how their schedulers and KV cache designs differ
• Demonstrated quantization work on real models (FP8/INT8/INT4) with measured accuracy impact and a defensible calibration methodology
• End-to-end experience with at least one of speculative decoding, distillation, or structured pruning — including the evaluation that validated it
• Strong Python; comfortable reading and patching PyTorch and serving-engine source code
• CUDA fundamentals: memory hierarchy, occupancy, kernel launch overhead, roofline reasoning; able to read and interpret a profiler trace
• Multi-GPU serving: tensor and pipeline parallelism, NCCL basics, and debugging distributed hangs and stragglers
• Kubernetes, Docker, and GPU scheduling in production; observability with Prometheus/Grafana or equivalent
• Ability to quantify your own work in latency percentiles, throughput per GPU, and cost per million tokens
Good to Have
• Real-time speech inference: streaming ASR/TTS, sub-second time-to-first-audio budgets, barge-in and turn-taking constraints
• Serving hybrid-architecture models (state-space/Mamba blocks combined with attention) and understanding their distinct state and cache management
• TensorRT-LLM engine building, NVIDIA NIM, or Triton Inference Server in production
• Experience with Indic or other multilingual and code-mixed models, and the evaluation discipline that comes with them
• Ray Serve, KServe, LLM gateways, or semantic and prefix cache layers
• Enough training-side exposure (Megatron, NeMo, DeepSpeed, FSDP) to work fluently with the pretraining team
• Public benchmarks, blog posts, or talks on inference optimization
What You Will Work On
You will work on in-house foundational models — not third-party API endpoints — running on large NVIDIA GPU clusters and serving live enterprise and government-scale voice traffic. The models are ours, so the whole stack is open to you: you can change the serving engine, the quantization recipe, the kernel, or the model itself, and you will often need to change more than one to hit a target.
The constraints are real and the feedback loop is fast. A latency regression is heard by callers within minutes. A successful optimization shows up directly in infrastructure spend. You will have access to proprietary speech and language data, in-house training and evaluation infrastructure, a close engineering partnership with NVIDIA, and a team that has been building Indian-language AI since 2016.
This role has a clear path to owning the inference platform and its technical direction, or to a deeper specialization in performance and kernel engineering, depending on where your strength lies.
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.