Live opening · Posted 2 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Responsibilities:
Lead SRE initiatives across all technology stacks.
Own the reliability, scalability, and performance of production systems.
Drive automation, observability, and incident response maturity.
Act as a technical authority across DevOps domains.
Architect and manage multi-cloud infrastructure (AWS, Azure) with IaC (Terraform, Helm).
Lead Kubernetes deployments across multi-cluster environments; optimize for scale and resilience.
Build and maintain CI/CD pipelines using GitLab, Jenkins, GitHub Actions; enforce GitOps workflows.
Implement and manage observability stack: Prometheus, Grafana, ELK, Splunk, Datadog.
Drive incident response, root cause analysis, and postmortem processes; reduce MTTR.
Enforce SRE principles: SLIs, SLOs, error budgets, chaos engineering.
Integrate DevSecOps practices: IAM, RBAC, secrets management, vulnerability scanning.
Enable MLOps workflows for AI/ML model deployment and lifecycle management.
Mentor junior SREs and DevOps engineers; establish technical standards and best practices.
Collaborate with product, engineering, and security teams to align reliability goals.
Own disaster recovery planning, backup strategies, and environment consistency.
Lead cost optimization and performance tuning across infrastructure layers.
Requirements:
10+ years in DevOps/SRE/Platform Engineering.
Deep expertise in Kubernetes, Docker, Terraform, Helm.
Strong proficiency in scripting (Python, Bash, Go).
Proven experience with cloud-native architectures and distributed systems.
Hands-on experience with CI/CD tooling and automation frameworks.
Familiarity with security frameworks and compliance requirements.
Demonstrated leadership in scaling systems and mentoring teams.
Strategic thinking across infrastructure and reliability domains.
Technical leadership with cross-functional influence.
High accountability and ownership of production systems.
Clear, concise communication across global teams.
Decision-making under pressure during incidents.
Ability to mentor and elevate team capabilities.
Bias for automation and continuous improvement.
Strong stakeholder management and alignment with business goals.
Experience
10-14 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.