Live opening · Posted 11 days ago

MLOps Engineer

Aivar Innovations · Bangalore | Coimbatore
Instahyre 3-7 yrs
You are 11 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 11 days ago
CompanyAivar Innovations
LocationBangalore | Coimbatore
Experience3-7 yrs
SourceInstahyre
Listed11 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
0 min from Instahyre publishing this role to us finding it
7 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
21,201 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

Own the JARK-Stack integration on EKS: Ray and KubeRay for distributed compute, Kubeflow Pipelines for workflow orchestration, MLflow for experiment tracking, JupyterHub for development, and advanced job schedulers (Kueue, Volcano, and Argo) for batch training. A bridge between data scientists and the platform.
Responsibilities:
Deploy and optimize Ray + KubeRay for distributed data processing and model training across GPU clusters.
Build Kubeflow Pipelines for reproducible ML workflows: data prep, training, evaluation, and deployment with lineage tracking.
Configure MLflow for centralized experiment tracking and model registry across teams.
Implement advanced job scheduling queue management, priority, preemption, and gang scheduling via Kueue/Volcano.
Build model CI/CD automated training, evaluation, validation, and canary/blue-green deployment to inference endpoints.
Create self-service tooling for data scientists' cluster provisioning, GPU allocation, and experiment templates.
Monitor ML workload performance, GPU utilization, training throughput, and data pipeline efficiency.
Requirements:
ML infrastructure / MLOps / ML platform engineering (3+ years).
Kubernetes (EKS preferred) deployments, PVs, RBAC, resource management.
At least two of: Ray/KubeRay, Kubeflow, MLflow, Airflow, Argo Workflows.
Distributed training with PyTorch DDP, Horovod, DeepSpeed, or Ray Train.
Model serving KServe, Seldon, or custom FastAPI serving.
GPU scheduling and resource management on Kubernetes.
Strong Python engineering tools and automation, not just notebooks.
Core Tech Stack: Ray/KubeRay, Kubeflow Pipelines, MLflow, JupyterHub, Argo Workflows, Kueue/Volcano, PyTorch/DeepSpeed, KServe, Helm, AWS (EKS, S3 EFS, ECR), Prometheus/Grafana.

Experience
3-7 yrs

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App