Live opening · Posted 27 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Requirements:
Bachelor's or Master's degree in Computer Science, Engineering, or a related field.
12+ years of experience in site reliability, production engineering, or large-scale distributed system operations.
Proven track record of designing and managing highly available, globally distributed systems in cloud-native environments (AWS, Azure, GCP).
Expert-level proficiency in one or more programming/scripting languages (Python, Go, Java, or Bash) for automation and tooling.
Deep understanding of Kubernetes, microservices, and service mesh architectures.
Advanced experience with Infrastructure as Code (Terraform, CloudFormation) and CI/CD automation frameworks.
Mastery in observability and monitoring stacks (Prometheus, Grafana, Datadog, OpenTelemetry).
Strong expertise in networking, storage, and distributed databases (both SQL and NoSQL).
Demonstrated ability to influence architectural decisions and drive reliability strategy across organizations.
Exceptional communication, leadership, and stakeholder management skills.
Experience
12-16 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.