Live opening · Posted 2 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
We are looking for an experienced SRE Engineer to ensure the reliability, scalability, and performance of large-scale distributed systems in a production environment.
Responsibilities:
Manage and optimise production environments across AWS, Kubernetes, Docker, and Linux/Unix.
Build and maintain CI/CD pipelines and Infrastructure as Code using Terraform and Ansible.
Implement observability and monitoring using Prometheus, Grafana, OpenTelemetry, or Datadog.
Set up SLO-based alerting and incident detection mechanisms.
Troubleshoot production issues and participate in incident response.
Improve system reliability, availability, scalability, and performance.
Participate in the production on-call rotation, including some weekends approximately once every 2-3 weeks.
Requirements:
4-6 years of experience in SRE, DevOps, or Production Engineering.
Strong hands-on experience with AWS, Kubernetes, Terraform, Grafana, and Prometheus.
Strong knowledge of Linux/Unix and Docker.
Experience with CI/CD and Infrastructure as Code.
Hands-on experience with observability, monitoring, alerting, and SLOs.
Experience working with large-scale distributed systems in production.
Experience
4-8 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.