Live opening · Posted 7 days ago

AI Systems & Inference Engineer (vLLM / TensorRT-LLM / Rust)

Quantum Tiger · Kolkata, West Bengal, India (On-site)
Linkedin No
You are 7 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 7 days ago
CompanyQuantum Tiger
LocationKolkata, West Bengal, India (On-site)
Work modeNo
SourceLinkedin
Listed7 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
4 min from Linkedin publishing this role to us finding it
17 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
39,799 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

If you know why standard cloud APIs fail in air-gapped sovereign environments, this role is for you.
When deploying high-concurrency LLMs inside enterprise and national infrastructure, sending API payloads over the public internet isn't an option. Neither is tolerating standard inference latency, unmanaged VRAM spikes, or unbounded KV cache fragmentation.
At Quantum Tiger, we are expanding our core systems engineering floor to build HyperEdge—our bare-metal compute engine and low-latency inference runtime powering sovereign AI deployments.
We are looking for an AI Systems & Inference Engineer who operates at the intersection of systems software, silicon performance, and model serving.
What You Will Build & Own:
Runtime Architecture: Engineer and optimize bare-metal LLM serving pipelines using vLLM, TensorRT-LLM, and Rust/C++.
Memory & Throughput Optimization: Implement continuous batching, PagedAttention tuning, and aggressive KV cache compression strategies to eliminate TTFT (Time-to-First-Token) bottlenecks.
Silicon Efficiency: Benchmark and deploy sub-4-bit quantization, FP8 runtime paths, and speculative decoding on private GPU clusters.
Air-Gapped Sovereign Deployment: Build deterministic, zero-data-leakage inference pipelines running entirely on-premises without external cloud dependency.
What We Expect:
Deep understanding of GPU memory hierarchy, CUDA execution models, and distributed inference topologies (Tensor/Pipeline parallelism).
Hands-on experience hacking on inference runtime engines (vLLM, TensorRT-LLM, TGI, or custom C++/Rust inference runtimes).
Instincts for low-level systems profiling (Nsight Systems, Perf, PyTorch Profiler).
A builder mindset: zero interest in building thin API wrappers, high appetite for solving hard compute problems from first principles.
We don’t believe in simulated work, bureaucracy, or multi-month onboarding queues. You get high context, hard architectural problems, and immediate ownership of sovereign systems that matter.
📍 Location: Kolkata (On-site / Hybrid)
📩 How to Apply: Send your GitHub, portfolio, or a teardown of an inference pipeline you optimized directly to future@quantumtiger.in/ manas@quantumtiger.in with the subject line AI Systems Engineer - [Your Name].
(Know an engineer who lives in the internals of model inference? Tag them below 👇)
#Hiring #SystemsEngineering #AIInfrastructure #vLLM #TensorRT #RustLang #DeepTech #QuantumTiger #SovereignAI #CUDA

Work arrangement
No

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App