Live opening · Posted 8 hours ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Requirements:
6+ years of strong hands-on software/platform engineering experience.
Strong programming skills in Python.
Deep hands-on experience with Kubernetes, containers, and Linux.
Experience building or operating large-scale distributed systems or platform infrastructure.
Strong understanding of cloud infrastructure, including compute, storage, and networking.
Experience building production-grade CI/CD and automation platforms.
Experience with workflow orchestration such as Apache Airflow or equivalent systems.
Experience with ML lifecycle tooling such as MLflow or equivalent platforms.
Strong understanding of ML training and deployment workflows.
Experience with Infrastructure as Code, preferably Terraform.
Strong debugging and production troubleshooting skills.
Experience building systems with monitoring, logging, metrics, and alerting.
Strong fundamentals in system design, reliability, and distributed computing.
Strong Differentiators:
We would especially love to meet you if you have worked on Ray or other distributed computing frameworks.
Distributed or multi-node ML training.
GPU and multi-GPU infrastructure.
Kubernetes-based ML platforms.
ML training platforms used by multiple teams.
Model serving and inference infrastructure.
GPU scheduling and utilization optimization.
Large-scale workflow orchestration.
ML platform developer experience.
Infrastructure cost and performance optimization.
Feature platforms or feature stores.
Good to Have:
Databricks, SageMaker, Vertex AI, or similar ML platforms.
Model monitoring, data drift, and automated retraining.
NVIDIA Triton, Ray Serve, vLLM, or SGLang.
LLM training/inference and LLMOps.
Vector databases and retrieval infrastructure.
Model governance and lineage.
OpenTelemetry, Grafana, or similar observability ecosystems.
Skills
Kubernetes, Linux, ML platform, ML platforms, MLOPS, Python, containers, ml ops
Experience
6-10 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.