Live opening · Posted 1 day ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
About the Company
Role - Cloud Reliability Engineer
Location - PAN India
About the Role
Job Requirements
Responsibilities
Define and maintain service SLIs, SLOs, error budgets, reliability dashboards, and operational scorecards.
Design, build, and operate secure, highly available, scalable, and cost-conscious cloud platforms and services.
Implement observability covering metrics, logs, traces, synthetic checks, business signals, alerting, and diagnostic context.
Improve alert quality, reduce noise, and establish actionable escalation and incident-response processes.
Lead or support incident response, root-cause analysis, corrective actions, and blameless postmortems.
Automate provisioning, deployment, configuration, diagnostics, remediation, maintenance, and recovery activities.
Implement self-healing, automated rollback, progressive delivery, canary, blue-green, and safe deployment practices.
Operate and optimize Kubernetes workloads, autoscaling, workload placement, resource limits, health checks, and availability controls.
Perform load, stress, soak, failover, chaos, and disaster-recovery testing to identify weaknesses before production impact.
Partner with development teams to embed reliability, observability, security, and operability into the software lifecycle.
Support cloud migration readiness, dependency assessment, baseline measurement, and post-migration reliability validation.
Conduct capacity planning, performance tuning, resource optimization, and FinOps-aligned cost improvement.
Maintain runbooks, architecture and support documentation, audit evidence, operational standards, and knowledge assets.
Mentor engineers and promote SRE, automation, continuous learning, and shared operational ownership.
Qualifications
6-12 years of experience in Cloud Engineering, Site Reliability Engineering, DevOps, Platform Engineering, or Production Engineering.
Hands-on experience operating production workloads on AWS, Azure, and/or GCP.
Strong knowledge of SLIs, SLOs, error budgets, availability, latency, throughput, saturation, and reliability risk.
Experience with Prometheus, Grafana, Dynatrace, CloudWatch, Azure Monitor, Google Cloud Monitoring, OpenTelemetry, ELK, Datadog, or comparable tools.
Hands-on experience with Docker, Kubernetes, and managed services such as EKS, AKS, or GKE.
Experience with Terraform or CloudFormation and reproducible environment provisioning.
Strong CI/CD experience using Jenkins, GitHub Actions, GitLab CI, Azure DevOps, or equivalent tools.
Proficiency in Python, Bash, Shell, PowerShell, or Go for automation, diagnostics, and operational tooling.
Knowledge of incident response, root-cause analysis, blameless postmortems, problem management, and on-call practices.
Experience with resilience patterns such as timeouts, retries, circuit breakers, bulkheads, graceful degradation, and automated rollback.
Knowledge of high availability, disaster recovery, capacity planning, performance testing, chaos engineering, and failover validation.
Understanding of IAM, least privilege, secrets, encryption, vulnerability remediation, auditability, and DevSecOps controls.
Required Skills
Cloud migration and modernization reliability validation.
Distributed microservices, API platforms, messaging, or high-volume transactional systems.
Preferred Skills
Cloud migration and modernization reliability validation.
Distributed microservices, API platforms, messaging, or high-volume transactional systems.
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.