Live opening · Posted 18 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Responsibilities:
ML Platform Architecture: Design and scale a distributed ML platform for data, feature store, training, and inferencing; build multi-region, containerized infrastructure using Kubernetes and Terraform.
Training and Inferencing at Scale: Optimize GPU/TPU/CPU utilization for large-scale training and real-time inferencing; implement distributed training, model parallelism, and caching with Ray.
Performance and Governance: Drive FinOps-aligned architecture, auto-scaling, and efficiency; enable observability, SLA/SLO tracking, and incident management.
Collaboration and Leadership: Partner with cross-functional teams to deliver high-impact ML solutions; mentor engineers and set best practices for ML systems design.
Requirements:
8-12 years in ML platform/distributed systems.
Strong in Python, Kubernetes, and GPU/TPU optimization.
Proven ability to design fault-tolerant, high-throughput ML systems.
Technology Stack:
ML Frameworks: TensorFlow, PyTorch.
Serving: Triton, TensorFlow Serving, TorchServe.
Distributed Compute: Ray, Kubernetes, Spark, and GPU/TPU optimization.
Lifecycle: MLflow and AirFlow.
Infra: AWS/GCP/Azure, Terraform, CI/CD.
Experience
6-10 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.