Live opening · Posted 27 days ago

Senior ML Platform Engineer (LLM)

eBay · Bangalore
Instahyre 5-9 yrs
You are 27 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 27 days ago
CompanyeBay
LocationBangalore
Experience5-9 yrs
SourceInstahyre
Listed27 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
0 min from Instahyre publishing this role to us finding it
9 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
21,224 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

As an LLM Inference Engineer on our AI Platform team, you'll remove the compute-scaling bottleneck for production LLMs. Your job is to make frontier-model inference fast, efficient, reliable, and observable, the last mile from GPUs to APIs that products depend on. This role sits at the intersection of HPC, GPU systems, and MLOps and requires strong intuition for how model architecture, runtimes, and hardware interact.
Responsibilities:
Own production inference: Take models from handoff to production-grade serving, including release engineering, capacity planning, cost optimization, and incident response.
Tune inference performance: reduce end-to-end latency and increase throughput across real production traffic patterns.
Optimize runtimes and servers: Scale inference across heterogeneous GPU fleets; optimize stacks such as vLLM, Triton, and related components (e. g., schedulers, KV cache, batching, and memory).
Benchmark and measure: Build benchmarking suites, metrics, and tooling to quantify latency, throughput, GPU utilization, memory, and cost.
Reliability and observability: Improve monitoring, tracing, and alerting; participate in incident response and postmortems to harden systems.
Apply and ship new optimizations: Evaluate research and implement pragmatic inference optimizations (e. g., quantization, paging, and kernel/runtime improvements).
Partner cross-functionally: Work with data science and product teams to translate business requirements into performance and availability SLOs.
Requirements:
Experience deploying and operating LLM inference services in production.
Strong production coding skills in Python plus Go or Rust (systems-level implementation and debugging).
Experience with ML frameworks and runtimes: PyTorch, vLLM, SGLang (and/or TensorRT).
Knowledge of GPU architecture and performance (profiling, memory bandwidth/latency tradeoffs); CUDA/kernel programming is a strong plus.
Solid understanding of LLM inference and optimization techniques: continuous batching, KV cache management, quantization, speculative decoding (nice-to-have), etc.
4-5+ years of hands-on experience in performance optimization and systems programming for AI/ML workloads.
Demonstrated ability to deliver measurable production improvements (e. g., 2X throughput, lower p95/p99 latency, reduced GPU cost).
Proven skill in root-cause analysis: finding bottlenecks across model, runtime, networking, and infrastructure.

Experience
5-9 yrs

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App