Live opening · Posted 9 hours ago

Machine Learning Engineer II

Nykaa · Bengaluru, Karnataka, India (On-site)
Linkedin No
JobBeeper subscribers received an alert for this role.

At a glance

The key details from the original listing.

Posted 9 hours ago
CompanyNykaa
LocationBengaluru, Karnataka, India (On-site)
Work modeNo
SkillsMachine Learning, Python, AWS, Docker, Redis, TensorFlow, PyTorch
SourceLinkedin
ListedPosted 9 hours ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
5 min from Linkedin publishing this role to us finding it
10 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
76,395 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

Senior Applied ML Engineering | Platform Co-Ownership | Production Excellence
About the role
Lead production ML engineering across Recommendations, Ads, Search, and new initiatives. Own model deployment end to end; co-own Inference Engine feature development with the platform team. Partner with Scientists, who own model behavior and business outcomes, and ensure production models meet agreed latency, reliability, quality, and cost goals.
Key responsibilities
Inference Engine development and model deployment
Co-own the Inference Engine roadmap and feature delivery with the platform team: translate Data Science needs into requirements, and contribute to design, implementation, testing, documentation, integration, reliability, and adoption.
Own model deployment end to end: register artifacts and metadata, version models, validate in staging, coordinate API integration, promote to production, configure defaults, run shadow tests or go-live, and manage monitoring and rollback.
Set readiness gates for dependencies, feature access, request and response contracts, compute and storage, and load-test latency, throughput, errors, and cost.
Build reusable real-time and batch serving patterns for feature retrieval, model composition, and pipeline orchestration.
GPU-based inference and deep learning training
Lead GPU-backed deep learning and language-model inference; tune runtimes, batching, concurrency, precision, VRAM, and CPU-to-GPU transfers for latency, throughput, quality, and cost.
Build GPU training and fine-tuning workflows with efficient data loading, mixed precision, multi-GPU execution, checkpointing, and capacity planning; support small language model adaptation and distillation.
Benchmark utilization, VRAM, p95/p99 latency, throughput, and H100-class versus current GPU options; inform model, capacity, and cost decisions.
Reliability and technical leadership
Own SLOs, dashboards, alerts, and runbooks for service health, features, models, latency, errors, and cost; lead incident response and root-cause follow-up.
Improve resilience and efficiency through capacity planning, autoscaling, caching, graceful degradation, safe rollbacks, and cost reviews.
Partner with Scientists, platform, and product engineering; review designs and code, mentor engineers, and maintain reusable libraries, standards, and launch guidance.
Evaluate practical LLM, RAG, and AI-assisted tools for experimentation or operations, with clear quality, privacy, and cost checks.
What you will bring
Strong Python and software engineering skills, with experience building APIs or distributed production services.
Strong hands-on expertise training, fine-tuning, and serving deep learning models on GPUs using PyTorch or TensorFlow; optimize VRAM, precision, latency, and throughput.
Experience owning production model deployments, including packaging, versioning, staged validation, monitoring, and rollback.
Working knowledge of feature stores, low-latency data access, caching, data contracts, and performance profiling.
Familiarity with AWS, Docker, CI/CD, MLflow, FastAPI or equivalent, Redis or DynamoDB, and observability tooling.
Technical leadership across teams, clear trade-off communication, operational ownership, and focus on reliability and cost.
Useful experience
Shared ML platforms, real-time feature stores, vector search, shadow deployments, model orchestration, SLM fine-tuning, or RAG systems.
How success is measured
Models are promoted safely, monitored in production, and have clear operational ownership.
Inference Engine features are adopted and reduce deployment friction across Data Science teams.
Serving and GPU workloads meet agreed quality, latency, reliability, and cost goals.
Repeatable operating practices reduce incidents and infrastructure waste.

Work arrangement
No

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App