Live opening · Posted 5 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Join a fast-growing GenAI startup as an SRE and take ownership of the reliability, scalability, and operational excellence of a platform powering the end-to-end ML lifecycle. You’ll work across Kubernetes, cloud infrastructure, production systems, observability, automation, and model-serving workloads, while helping build strong SRE and incident-management practices.
What You'll Do
Own platform uptime, reliability, scalability, and performance
Manage Kubernetes clusters, cloud infrastructure, and production environments
Establish and improve incident response, on-call, RCA, and postmortem processes
Drive deployment, rollback, and change-management practices
Handle capacity planning and disaster recovery
Build and enhance monitoring, alerting, and operational dashboards
Automate deployments, scaling, and repetitive operational workflows
Troubleshoot complex production infrastructure and application issues
Support GPU workloads and model-serving infrastructure
Must have
Experience in SRE, DevOps, or Platform Engineering Strong hands-on knowledge of Linux, networking, and Kubernetes Experience with AWS, GCP, or Azure Hands-on experience with Terraform, Helm, and CI/CD Strong troubleshooting, incident management, and problem-solving skills Scripting/programming experience in Python, Bash, or Go
Good to have
Experience with MLOps / AI infrastructure Exposure to GPU clusters and model serving Knowledge of release engineering Understanding of reliability and operational best practices
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.