Live opening · Posted 6 hours ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Requirements:
6+ years of strong hands-on software/platform engineering experience.
Strong programming skills in Python.
Deep hands-on experience with Kubernetes, containers and Linux.
Experience building or operating large-scale distributed systems or platform infrastructure.
Strong understanding of cloud infrastructure including compute, storage and networking.
Experience building production-grade CI/CD and automation platforms.
Experience with workflow orchestration such as Apache Airflow or equivalent systems.
Experience with ML lifecycle tooling such as MLflow or equivalent platforms.
Strong understanding of ML training and deployment workflows.
Experience with Infrastructure as Code, preferably Terraform.
Strong debugging and production troubleshooting skills.
Experience building systems with monitoring, logging, metrics and alerting.
Strong fundamentals in system design, reliability and distributed computing.
Strong Differentiators:
We would especially love to meet you if you have worked on:
Ray or other distributed computing frameworks
Distributed or multi-node ML training
GPU and multi-GPU infrastructure
Kubernetes-based ML platforms
ML training platforms used by multiple teams
Model serving and inference infrastructure
GPU scheduling and utilisation optimisation
Large-scale workflow orchestration
ML platform developer experience
Infrastructure cost and performance optimisation
Feature platforms or feature stores
Skills
Airflow, CI/CD, Docker, Kubernetes, Linux, MLOps, MLflow, Python, Terraform
Experience
6-9 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.