Live opening · Posted 26 days ago

Machine Learning Engineer 4

Adobe · Bangalore | Noida
Instahyre 7-11 yrs
You are 26 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 26 days ago
CompanyAdobe
LocationBangalore | Noida
Experience7-11 yrs
SourceInstahyre
Listed26 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
0 min from Instahyre publishing this role to us finding it
16 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
21,186 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

Join a team where ambition meets brand new ideas! Adobe's Firefly Group offers an outstanding opportunity to be at the forefront of innovative technology. As leaders in digital experiences, we are determined to transform how companies interact with customers across every screen. Our mission is to empower engineers and researchers by providing world-class cloud solutions for ML workflows, using the latest technology and frameworks. If you are passionate about crafting flawless digital experiences and are eager to compete in a dynamic environment, this is the place for you.
The AI Platform team builds the compute infrastructure that enables AI workloads at scale, spanning job scheduling, resource management, and the control plane that ASML engineers at Adobe depend on daily. We sit at the intersection of distributed systems and machine learning infrastructure, building the foundational services that power model training across Adobe Firefly.
We are seeking a seasoned Senior ML Platform engineer to join our ambitious team in Noida/Bangalore. This role is ideal for a technical leader who can own and evolve complex scheduling and compute orchestration systems, work closely with ASML engineers as the primary consumers of the platform, and drive architectural decisions that meaningfully improve developer productivity and system reliability.
Responsibilities:
Design, build, and maintain core services of the AI compute control plane: job scheduling, cluster management, resource quota enforcement, and compute lifecycle management.
Lead the design and implementation of job scheduling, resource quota enforcement, and compute lifecycle management systems.
Own the control plane services that manage GPU/CPU workload orchestration from job submission through execution, monitoring, and teardown.
Design reliable, fault-tolerant worker services and supervisor patterns for long-running compute workloads.
Build and evolve the data layer that tracks job state, cluster state, and resource ownership across the platform.
Partner closely with ASML engineers to deeply understand their workflows and translate requirements into robust platform capabilities.
Develop and maintain Python SDKs and CLIs that ML engineers use to interact with the platform, prioritising developer experience and reliability.
Drive end-to-end ownership of features from API design and data modelling through deployment and production operations.
Establish observability standards (metrics, tracing, alerting) for scheduling and compute systems.
Lead incident response and root cause analysis for production issues in compute orchestration.
Mentor junior and mid-level engineers on system design, scheduling patterns, and platform engineering best practices.
Requirements:
B. Tech / M. Tech degree in Computer Science from a premier institute.
9+ years of proven experience in backend platform engineering, distributed systems, or infrastructure software.
Strong computer science fundamentals, particularly in distributed systems, concurrency, and system design.
Experience building or operating job scheduling, workflow orchestration, or compute management systems (e. g., Argo, Airflow, Ray, Slurm, or similar).
Proficiency in Python and/or Java, with strong async programming skills.
Experience designing and operating services backed by relational databases (PostgreSQL preferred) at scale.
Deep understanding of Cloud Platforms, with preference for AWS; familiarity with Azure or GCP is a plus.
Proven track record of working directly with internal engineering customers (ML engineers, researchers) to shape platform roadmap.
Strong problem-solving skills with the ability to own ambiguous, complex systems independently.
Experience with Kubernetes at the workload/scheduling layer (not just operations).
Good to Have:
Hands-on experience with ML training workflows, distributed training frameworks (PyTorch, TensorFlow), or GPU resource management.
Familiarity with gRPC/protobuf or event-driven architectures.
Experience building developer-facing internal platforms consumed by ML or research teams.
Prior work in an AI platform, MLOps, or compute infrastructure role.
Understanding of ML lifecycle experiment tracking, model versioning, and training pipelines.

Experience
7-11 yrs

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App