Live opening · Posted 3 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
About The Role
The Site Reliability Engineer owns the systems and practices that keep production services available, performant, and recoverable at scale. The role focuses on Kubernetes-based workloads, cloud infrastructure, observability, incident response, and automation across distributed systems.
This engineer will work with platform, software, and security teams to reduce operational risk and improve the developer experience. The work includes defining service-level objectives, eliminating recurring failure modes, and building reliable delivery and recovery processes for customer-facing systems.
Key Responsibilities
Design and operate highly available infrastructure on AWS or GCP using Kubernetes, Terraform, and infrastructure-as-code standards
Build and maintain observability across services using Prometheus, Grafana, OpenTelemetry, and centralized logging platforms such as ELK or Datadog
Define and track SLIs, SLOs, error budgets, and capacity plans for critical production services
Automate deployment, scaling, backup, and recovery workflows through CI/CD pipelines using tools such as GitHub Actions, GitLab CI, or Argo CD
Lead incident response, including triage, communications, root-cause analysis, and durable remediation of recurring production issues
Harden systems through access controls, secrets management, patching, disaster recovery testing, and infrastructure reliability reviews
Partner with application teams to improve service design, operational readiness, performance, and on-call effectiveness
What We Are Looking For
3–8 years of experience in site reliability engineering, DevOps, platform engineering, or production infrastructure roles
Hands-on experience operating Kubernetes and containerized services in AWS, GCP, or Azure
Strong Linux administration and networking fundamentals, including DNS, HTTP, TLS, TCP/IP, load balancing, and troubleshooting distributed systems
Proficiency with Terraform or an equivalent infrastructure-as-code tool, plus practical experience designing CI/CD pipelines
Experience implementing observability with metrics, logs, traces, dashboards, and actionable alerting using tools such as Prometheus, Grafana, OpenTelemetry, or Datadog
Strong scripting or programming skills in Python, Go, or Bash, with a focus on automation, testing, and maintainable operational tooling
Bonus: Experience with service meshes, Kafka, PostgreSQL or Redis operations, compliance-focused infrastructure, chaos engineering, or a bachelor's degree in computer science, engineering, or a related field
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.