Live opening · Posted 8 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
About The Role
The Site Reliability Engineer will design, operate, and improve the production infrastructure supporting high-volume, customer-facing services. The role focuses on Kubernetes, AWS, observability, incident response, and automation across distributed systems where availability, latency, and operational safety are critical.
The engineer will partner with software teams to define service-level objectives, eliminate recurring failure modes, and build reliable deployment and recovery workflows. This role has direct ownership of production health and meaningful influence over platform architecture, engineering standards, and operational practices.
Key Responsibilities
Operate and improve highly available production services running on AWS and Kubernetes, including capacity planning, scaling, and failure recovery
Define and maintain SLIs, SLOs, error budgets, and operational dashboards using tools such as Prometheus, Grafana, and OpenTelemetry
Automate infrastructure provisioning and configuration with Terraform, Helm, and GitHub Actions or equivalent CI/CD systems
Lead incident response, coordinate technical mitigation, and produce clear post-incident reviews with measurable corrective actions
Harden deployment pipelines with progressive delivery, automated rollback, health checks, and change-management controls
Identify and eliminate recurring toil through Python or Go automation, platform tooling, and self-service workflows
Collaborate with application engineers on performance tuning, resilience testing, dependency management, and production readiness reviews
What We Are Looking For
3–8 years of experience in site reliability engineering, DevOps, platform engineering, or a closely related production infrastructure role
Hands-on experience operating Kubernetes workloads in production, including deployments, networking, storage, ingress, and troubleshooting
Strong knowledge of AWS services such as EC2, EKS, IAM, VPC, RDS, S3, and CloudWatch
Proficiency with infrastructure as code and delivery tooling, including Terraform, Helm, Git, and CI/CD pipelines
Experience building observability systems with metrics, logs, traces, alerting, and on-call practices using tools such as Prometheus, Grafana, Datadog, or OpenTelemetry
Proficiency in Python, Go, or a similar programming language, with a track record of replacing manual operational work with reliable automation
Bachelor’s degree in computer science, engineering, or a related technical field, or equivalent practical experience
Bonus: experience with service meshes, distributed systems, chaos engineering, compliance-focused infrastructure, or multi-region production environments
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.