Live opening · Posted 7 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
We are hiring a hands-on Site Reliability Engineer (SRE) to support the reliability, availability, security, and performance of cloud-native production systems. This role works closely with engineering, product, and support teams and requires a strong ownership mindset. You will be involved in building, operating, and improving services at scale, with a strong focus automation, observability, and operational excellence This role operates exclusively on the night shift from India to support US-based teams and on-call rotations.
Responsibilities:
Operate and support mission-critical production systems on AWS with high availability and reliability.
Build, maintain, and enhance CI/CD pipelines for secure and repeatable deployments.
Deploy and manage Kubernetes workloads using Docker and Helm charts.
Implement and operate monitoring, logging, and alerting using Prometheus, Grafana, and Loki.
Automate infrastructure and operational workflows using Python and Terraform.
Manage infrastructure and secrets using HashiCorp tools (Terraform, Vault).
Participate in incident response, escalation handling, RCA, and post-incident improvements.
Collaborate closely with engineering, product, and support teams during US business hours.
Continuously improve system reliability, security posture, and operational efficiency.
Requirements:
3-6 years of hands-on experience in SRE, DevOps, or infrastructure engineering roles.
Strong experience supporting production systems at scale.
Hands-on AWS experience (EC2 VPC, IAM, EKS, S3 etc. ).
Solid experience with CI/CD tools (GitHub Actions, GitLab CI, Jenkins, AWS CodePipeline, etc. ).
Hands-on experience with Kubernetes and Docker, including Helm-based deployments.
Strong experience with Prometheus-based observability (Grafana, Loki).
Proficiency in Python or Bash scripting for automation.
Strong experience with Infrastructure as Code using Terraform.
Exposure to HashiCorp Vault or secrets management best practices.
Good understanding of security, compliance, and cloud-native operations.
Strong written and verbal communication skills.
Willingness to work the night shift and participate in on-call rotations supporting US teams.
Nice to Have:
Experience with SLIs, SLOs, and error budgets.
Exposure to multi-cloud (GCP / Azure).
Prior experience in SaaS, security, or high-availability platforms.
Experience
2-6 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.