Live opening · Posted 7 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
If you know why standard cloud APIs fail in air-gapped sovereign environments, this role is for you.
When deploying high-concurrency LLMs inside enterprise and national infrastructure, sending API payloads over the public internet isn't an option. Neither is tolerating standard inference latency, unmanaged VRAM spikes, or unbounded KV cache fragmentation.
At Quantum Tiger, we are expanding our core systems engineering floor to build HyperEdge—our bare-metal compute engine and low-latency inference runtime powering sovereign AI deployments.
We are looking for an AI Systems & Inference Engineer who operates at the intersection of systems software, silicon performance, and model serving.
What You Will Build & Own:
Runtime Architecture: Engineer and optimize bare-metal LLM serving pipelines using vLLM, TensorRT-LLM, and Rust/C++.
Memory & Throughput Optimization: Implement continuous batching, PagedAttention tuning, and aggressive KV cache compression strategies to eliminate TTFT (Time-to-First-Token) bottlenecks.
Silicon Efficiency: Benchmark and deploy sub-4-bit quantization, FP8 runtime paths, and speculative decoding on private GPU clusters.
Air-Gapped Sovereign Deployment: Build deterministic, zero-data-leakage inference pipelines running entirely on-premises without external cloud dependency.
What We Expect:
Deep understanding of GPU memory hierarchy, CUDA execution models, and distributed inference topologies (Tensor/Pipeline parallelism).
Hands-on experience hacking on inference runtime engines (vLLM, TensorRT-LLM, TGI, or custom C++/Rust inference runtimes).
Instincts for low-level systems profiling (Nsight Systems, Perf, PyTorch Profiler).
A builder mindset: zero interest in building thin API wrappers, high appetite for solving hard compute problems from first principles.
We don’t believe in simulated work, bureaucracy, or multi-month onboarding queues. You get high context, hard architectural problems, and immediate ownership of sovereign systems that matter.
📍 Location: Kolkata (On-site / Hybrid)
📩 How to Apply: Send your GitHub, portfolio, or a teardown of an inference pipeline you optimized directly to future@quantumtiger.in/ manas@quantumtiger.in with the subject line AI Systems Engineer - [Your Name].
(Know an engineer who lives in the internals of model inference? Tag them below 👇)
#Hiring #SystemsEngineering #AIInfrastructure #vLLM #TensorRT #RustLang #DeepTech #QuantumTiger #SovereignAI #CUDA
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.