Live opening · Posted 9 hours ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Senior Applied ML Engineering | Platform Co-Ownership | Production Excellence
About the role
Lead production ML engineering across Recommendations, Ads, Search, and new initiatives. Own model deployment end to end; co-own Inference Engine feature development with the platform team. Partner with Scientists, who own model behavior and business outcomes, and ensure production models meet agreed latency, reliability, quality, and cost goals.
Key responsibilities
Inference Engine development and model deployment
Co-own the Inference Engine roadmap and feature delivery with the platform team: translate Data Science needs into requirements, and contribute to design, implementation, testing, documentation, integration, reliability, and adoption.
Own model deployment end to end: register artifacts and metadata, version models, validate in staging, coordinate API integration, promote to production, configure defaults, run shadow tests or go-live, and manage monitoring and rollback.
Set readiness gates for dependencies, feature access, request and response contracts, compute and storage, and load-test latency, throughput, errors, and cost.
Build reusable real-time and batch serving patterns for feature retrieval, model composition, and pipeline orchestration.
GPU-based inference and deep learning training
Lead GPU-backed deep learning and language-model inference; tune runtimes, batching, concurrency, precision, VRAM, and CPU-to-GPU transfers for latency, throughput, quality, and cost.
Build GPU training and fine-tuning workflows with efficient data loading, mixed precision, multi-GPU execution, checkpointing, and capacity planning; support small language model adaptation and distillation.
Benchmark utilization, VRAM, p95/p99 latency, throughput, and H100-class versus current GPU options; inform model, capacity, and cost decisions.
Reliability and technical leadership
Own SLOs, dashboards, alerts, and runbooks for service health, features, models, latency, errors, and cost; lead incident response and root-cause follow-up.
Improve resilience and efficiency through capacity planning, autoscaling, caching, graceful degradation, safe rollbacks, and cost reviews.
Partner with Scientists, platform, and product engineering; review designs and code, mentor engineers, and maintain reusable libraries, standards, and launch guidance.
Evaluate practical LLM, RAG, and AI-assisted tools for experimentation or operations, with clear quality, privacy, and cost checks.
What you will bring
Strong Python and software engineering skills, with experience building APIs or distributed production services.
Strong hands-on expertise training, fine-tuning, and serving deep learning models on GPUs using PyTorch or TensorFlow; optimize VRAM, precision, latency, and throughput.
Experience owning production model deployments, including packaging, versioning, staged validation, monitoring, and rollback.
Working knowledge of feature stores, low-latency data access, caching, data contracts, and performance profiling.
Familiarity with AWS, Docker, CI/CD, MLflow, FastAPI or equivalent, Redis or DynamoDB, and observability tooling.
Technical leadership across teams, clear trade-off communication, operational ownership, and focus on reliability and cost.
Useful experience
Shared ML platforms, real-time feature stores, vector search, shadow deployments, model orchestration, SLM fine-tuning, or RAG systems.
How success is measured
Models are promoted safely, monitored in production, and have clear operational ownership.
Inference Engine features are adopted and reduce deployment friction across Data Science teams.
Serving and GPU workloads meet agreed quality, latency, reliability, and cost goals.
Repeatable operating practices reduce incidents and infrastructure waste.
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.