Live opening · Posted 8 hours ago

AI Inference Platform Engineer

DRW · Chicago
Greenhouse
You are 8 hours behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 8 hours ago
CompanyDRW
LocationChicago
SourceGreenhouse
Listed8 hours ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
24 min from Greenhouse publishing this role to us finding it
6 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
71,275 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

DRW is a diversified trading firm with over 3 decades of experience bringing sophisticated technology and exceptional people together to operate in markets around the world. We value autonomy and the ability to quickly pivot to capture opportunities, so we operate using our own capital and trading at our own risk.
Headquartered in Chicago with offices throughout the U.S., Canada, Europe, and Asia, we trade a variety of asset classes including Fixed Income, ETFs, Equities, FX, Commodities and Energy across all major global markets. We have also leveraged our expertise and technology to expand into three non-traditional strategies: real estate, venture capital and cryptoassets.
We operate with respect, curiosity and open minds. The people who thrive here share our belief that it’s not just what we do that matters–it's how we do it. DRW is a place of high expectations, integrity, innovation and a willingness to challenge consensus.
About the Role
We're looking for an AI Inference Platform Engineer to build, operate, and optimize the systems that serve large language, vision, multimodal, and embedding models across DRW. This role provides DRW's firmwide interface to modern AI models, from early evaluation through reliable production use.
You'll work across inference runtimes, distributed systems, and production platform engineering, with deep GPU literacy. You'll own the serving platform end-to-end: onboarding newly released models, measuring quality and performance equivalence across serving configurations, scheduling workloads across tenants, and continuously improving latency, throughput, utilization, reliability, and cost across the inference fleet.
What You'll Do
Optimize LLM inference performance across modern NVIDIA GPU architectures and inference runtimes.
Build end-to-end performance profiling and observability to identify bottlenecks from individual GPU kernels through multi-node inference systems.
Design and optimize KV cache and distributed inference architectures, including caching, routing, memory tiering, and prefill/decode strategies.
Own day-0 model onboarding, determining the appropriate runtime, precision, sharding, memory, batching, cache policy, and serving configuration for new models.
Maintain validated performance profiles for important model and hardware combinations, including performance and quality regression testing.
Measure and monitor quality equivalence across serving configurations, including KV cache quantization, speculative decoding acceptance thresholds, precision choices, and model routing, so in-house serving can be trusted to match reference-model quality on production workloads.
Manage the production serving lifecycle of models, including versioning, compatibility, staging, canarying, promotion, rollback, and retirement.
Partner with SRE and platform teams to automate model deployment, distribution, production readiness, observability, and reliable operation across environments.
Optimize model placement, scaling, and resource allocation across the inference fleet to improve utilization and cost efficiency while meeting performance and reliability requirements.
Design and operate multi-tenant scheduling and isolation across shared GPU capacity, balancing latency SLOs, throughput, and priority across concurrent workloads.
What We're Looking For
The Tech
Hands-on experience serving LLMs on NVIDIA GPUs, with familiarity across current and emerging architectures (Hopper, Blackwell, and successors), HBM, Tensor Cores, NVLink/NVSwitch, and the compute and memory bottlenecks that shape serving decisions.
Deep expertise in at least one modern inference runtime such as TensorRT-LLM, vLLM, or SGLang.
Practical knowledge of inference optimization techniques including continuous batching, scheduling, chunked prefill, speculative decoding, quantization, CUDA Graphs, and paged attention.
Understanding of KV cache architecture, including prefix caching, block management, sizing, eviction, quantization, cache-aware routing, and multi-tier caching.
Experience measuring model quality equivalence across serving configurations, including evaluation harnesses, task-specific benchmarks, and regression detection for quantization, KV cache, and speculative decoding changes.
Experience designing and tuning distributed inference systems, including tensor parallelism, multi-node deployments, and disaggregated prefill and decode.
Experience with multi-tenant GPU scheduling, workload isolation, and QoS across concurrent inference workloads.
Proficiency with GPU performance and observability tooling such as Nsight, DCGM, OpenTelemetry, Prometheus, and Grafana.

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App