Live opening · Posted 9 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
We are hiring a Site Reliability Engineer (SRE) to help build and operate reliable, secure, and scalable cloud platforms on AWS. This is a hands-on platform and reliability engineering role. The focus is on AWS infrastructure, automation, observability, CI/CD, and operational reliability, supporting production applications and platform services.
Responsibilities:
Build, manage, and operate AWS infrastructure for production workloads.
Automate cloud infrastructure using Terraform and/or Pulumi.
Implement monitoring, alerting, logging, and operational dashboards.
Improve platform reliability, availability, performance, and scalability.
Build and maintain CI/CD pipelines and deployment automation.
Support containerized workloads running on ECS, EKS, or Kubernetes.
Troubleshoot infrastructure, networking, deployment, and production issues.
Implement operational practices around incident management, backups, recovery, and capacity.
Work with engineering teams to define reusable infrastructure and deployment patterns.
Improve developer experience through automation, self-service capabilities, and platform tooling.
Apply AWS security and IAM best practices across platform components.
Requirements:
We are looking for someone comfortable working hands-on with AWS and Infrastructure as Code, enjoys automating repetitive operational work, and can help engineering teams run production services reliably.
Strong AWS competency and practical Terraform/Pulumi experience are core expectations for this role.
1-3 years of experience in SRE, DevOps, cloud engineering, platform engineering, or a similar role.
Strong hands-on competency with AWS.
Experience with AWS services such as IAM, VPC and networking, EC2 ECS/EKS, Lambda, API Gateway, S3 CloudWatch, DynamoDB, or RDS.
Hands-on experience with Terraform and/or Pulumi for Infrastructure as Code.
Good understanding of Docker and containerized applications.
Experience with CI/CD platforms such as GitLab CI, GitHub Actions, or Jenkins.
Good understanding of Linux, networking, DNS, HTTP, and basic cloud security.
Ability to troubleshoot production systems using logs, metrics, and traces.
Understanding of reliability concepts such as availability, monitoring, alerting, retries, scaling, and disaster recovery.
Basic scripting/programming ability in Python, Bash, or a similar language.
Good to Have:
Experience building or operating an Internal Developer Platform (IDP).
Experience with Backstage or other developer portal technologies.
Kubernetes and EKS experience.
Experience creating reusable Terraform/Pulumi modules and platform templates.
OpenTelemetry, Prometheus, Grafana, or similar observability technologies.
Experience with event-driven architectures and AWS services such as SQS, SNS, and EventBridge.
Exposure to AWS Well-Architected practices, security controls, and cost optimization.
Experience supporting AI/ML, data, or agentic workloads on AWS.
Understanding of SRE practices such as SLIs, SLOs, error budgets, incident response, and post-incident reviews.
Experience
1-3 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.