Live opening · Posted 9 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Cloud Reliability Engineer (SRE)
Experience: 6-12 Years
Location: PAN India
Employment Type: Full-Time
Job Summary
We are looking for an experienced Cloud Reliability Engineer with strong expertise in Site Reliability Engineering (SRE), Cloud Platforms, Kubernetes, Observability, Automation, and DevOps. The role involves ensuring high availability, reliability, scalability, performance, and operational excellence of enterprise cloud platforms.
Mandatory Skills
6-12 years of experience in SRE, DevOps, Cloud Engineering, Platform Engineering, or Production Engineering.
Hands-on experience with AWS and/or Azure and/or GCP production environments.
Strong understanding of SLI, SLO, Error Budgets, Incident Management, RCA, and On-Call Operations.
Experience with Prometheus, Grafana, CloudWatch, Dynatrace, Datadog, ELK, OpenTelemetry, or similar monitoring tools.
Expertise in Docker, Kubernetes, Helm, and EKS/AKS/GKE.
Hands-on experience with Terraform and/or CloudFormation.
Strong CI/CD experience using Jenkins, GitHub Actions, GitLab CI, Azure DevOps, and GitOps practices.
Proficiency in Python, Shell/Bash, PowerShell, or Go.
Experience in High Availability, Disaster Recovery, Capacity Planning, Performance Testing, Failover Testing, and Chaos Engineering.
Knowledge of IAM, Security Controls, Secrets Management, Encryption, and DevSecOps.
Good to Have
Istio / Service Mesh
ArgoCD / Flux
Karpenter
Kafka
Splunk or Advanced Datadog
FinOps
Policy as Code
Application Performance Engineering
Cloud Migration & Modernization Projects
Distributed Microservices and API Platforms
Key Responsibilities
Build and operate highly available, secure, and scalable cloud platforms.
Implement observability, monitoring, alerting, and self-healing solutions.
Drive incident response, troubleshooting, root cause analysis, and postmortems.
Automate infrastructure provisioning, deployments, and operational processes.
Optimize Kubernetes workloads for performance, scalability, and reliability.
Collaborate with engineering teams to improve reliability, security, and operational excellence.
Preferred Certifications
AWS Solutions Architect (Associate/Professional)
AWS DevOps Engineer Professional
Azure DevOps Engineer Expert
Google Cloud Professional Cloud DevOps Engineer
Certified Kubernetes Administrator (CKA)
HashiCorp Terraform Associate
Cloud Reliability Engineer | SRE | Kubernetes | AWS/Azure/GCP | Terraform | DevOps | Observability
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.