Live opening · Posted 26 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Join a team where ambition meets brand new ideas! Adobe's Firefly Group offers an outstanding opportunity to be at the forefront of innovative technology. As leaders in digital experiences, we are determined to transform how companies interact with customers across every screen. Our mission is to empower engineers and researchers by providing world-class cloud solutions for ML workflows, using the latest technology and frameworks. If you are passionate about crafting flawless digital experiences and are eager to compete in a dynamic environment, this is the place for you.
The AI Platform team builds the compute infrastructure that enables AI workloads at scale, spanning job scheduling, resource management, and the control plane that ASML engineers at Adobe depend on daily. We sit at the intersection of distributed systems and machine learning infrastructure, building the foundational services that power model training across Adobe Firefly.
We are seeking a seasoned Senior ML Platform engineer to join our ambitious team in Noida/Bangalore. This role is ideal for a technical leader who can own and evolve complex scheduling and compute orchestration systems, work closely with ASML engineers as the primary consumers of the platform, and drive architectural decisions that meaningfully improve developer productivity and system reliability.
Responsibilities:
Design, build, and maintain core services of the AI compute control plane: job scheduling, cluster management, resource quota enforcement, and compute lifecycle management.
Lead the design and implementation of job scheduling, resource quota enforcement, and compute lifecycle management systems.
Own the control plane services that manage GPU/CPU workload orchestration from job submission through execution, monitoring, and teardown.
Design reliable, fault-tolerant worker services and supervisor patterns for long-running compute workloads.
Build and evolve the data layer that tracks job state, cluster state, and resource ownership across the platform.
Partner closely with ASML engineers to deeply understand their workflows and translate requirements into robust platform capabilities.
Develop and maintain Python SDKs and CLIs that ML engineers use to interact with the platform, prioritising developer experience and reliability.
Drive end-to-end ownership of features from API design and data modelling through deployment and production operations.
Establish observability standards (metrics, tracing, alerting) for scheduling and compute systems.
Lead incident response and root cause analysis for production issues in compute orchestration.
Mentor junior and mid-level engineers on system design, scheduling patterns, and platform engineering best practices.
Requirements:
B. Tech / M. Tech degree in Computer Science from a premier institute.
9+ years of proven experience in backend platform engineering, distributed systems, or infrastructure software.
Strong computer science fundamentals, particularly in distributed systems, concurrency, and system design.
Experience building or operating job scheduling, workflow orchestration, or compute management systems (e. g., Argo, Airflow, Ray, Slurm, or similar).
Proficiency in Python and/or Java, with strong async programming skills.
Experience designing and operating services backed by relational databases (PostgreSQL preferred) at scale.
Deep understanding of Cloud Platforms, with preference for AWS; familiarity with Azure or GCP is a plus.
Proven track record of working directly with internal engineering customers (ML engineers, researchers) to shape platform roadmap.
Strong problem-solving skills with the ability to own ambiguous, complex systems independently.
Experience with Kubernetes at the workload/scheduling layer (not just operations).
Good to Have:
Hands-on experience with ML training workflows, distributed training frameworks (PyTorch, TensorFlow), or GPU resource management.
Familiarity with gRPC/protobuf or event-driven architectures.
Experience building developer-facing internal platforms consumed by ML or research teams.
Prior work in an AI platform, MLOps, or compute infrastructure role.
Understanding of ML lifecycle experiment tracking, model versioning, and training pipelines.
Experience
7-11 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.