Live opening · Posted 6 hours ago

Senior Site Reliability Engineer (SRE) Engineer

Umanist NA · Pune City, Maharashtra, India (On-site)
Linkedin No
You are 6 hours behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 6 hours ago
CompanyUmanist NA
LocationPune City, Maharashtra, India (On-site)
Salary2M INR/yr - 2.5M INR/yr
Work modeNo
SkillsPython, AWS, Azure, GCP, Kubernetes, Terraform
SourceLinkedin
Listed6 hours ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
11 min from Linkedin publishing this role to us finding it
10 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
74,243 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

Senior Site Reliability Engineer (SRE) Engineer
ONLY PUNE, NAGPUR,KOLHAPUR PROFILES WILL BE CONSIDERED FOR INTERVIEW, A BIG "NO" FOR ANY OTHER LOCATIONS EVEN FOR MUMBAI.
"Microsoft Azure/AWS/GCP, Kubernetes, Terraform, Datadog, OpenTelemetry, Golden Signals monitoring(Latency,Traffic, Errors, Saturation) and modern SRE practices(SLIs, SLOs, SLAs, and Error Budgets.)" Please match your work experience with these mentioned skillset for a quick right-fit check.
Location: Viman Nagar, Pune – Work From Office
Experience Overall(must have): 8 Years
CTC: Up to ₹25 LPA
Notice Period: Immediate Joiners Only within 15d or (if serving max 30days)
Working Hours: 3:00 PM – 12:00 AM, Monday to Friday
On-Call: 24/7 Production Support – On-Call Rotation Required
Employment Type: Full-Time
About The Role
We are looking for an experienced Senior Site Reliability Engineer (SRE) / DevOps Engineer to manage and improve the reliability, scalability, performance, security, and observability of mission-critical production environments.
The role requires strong hands-on expertise in Cloud, Kubernetes, DevOps automation, Monitoring & Observability, Incident Management, and SRE practices. The ideal candidate should be comfortable handling production incidents while also driving long-term initiatives around reliability, automation, scalability, and reduction of operational toil.
Must-Have Skills & Experience1. SRE & Production Operations
Relevant 7+ years of relevant experience in SRE / DevOps / Cloud Infrastructure / Production Engineering.
Hands-on experience with 24/7 production support and on-call operations.
Strong experience in incident management, troubleshooting, RCA, and post-mortems.
Good understanding of SLI, SLO, SLA, Error Budgets, MTTR, and reliability engineering.
Experience with toil reduction, capacity planning, high availability, disaster recovery, and failover strategies.
Ability to improve system availability, performance, scalability, and operational reliability.
Cloud & Infrastructure
Strong hands-on experience with Microsoft Azure, AWS, and/or GCP.
Strong understanding of cloud infrastructure, networking, IAM, storage, compute, and cloud-native services.
Hands-on experience with at least one major cloud platform and good exposure to multi-cloud environments.
Experience with:
Azure: VMs, Networking, Storage, IAM, Azure Monitor, AKS
AWS: EC2, S3, RDS, IAM, VPC, CloudWatch, EKS
GCP: Compute Engine, Cloud Storage, IAM, VPC, GKE, Cloud Monitoring
Kubernetes & Containerization
Strong hands-on experience with Kubernetes and containerized workloads.
Experience with AKS / EKS / GKE or equivalent Kubernetes environments.
Hands-on experience with Helm deployments.
Understanding of Kubernetes troubleshooting, scaling, networking, and workload management.
Infrastructure as Code & DevOps
Hands-on experience with Terraform / Infrastructure as Code (IaC).
Experience with Git-based workflows using GitHub, GitLab, or Azure Repos.
Strong DevOps automation and CI/CD understanding.
Strong scripting skills in Python and/or Bash.
Monitoring & Observability
Strong hands-on experience with OpenTelemetry.
Experience with monitoring and observability tools such as:
Prometheus
Grafana
Datadog
Azure Monitor
AWS CloudWatch
GCP Cloud Monitoring
Strong understanding of metrics, logs, distributed tracing, and alerting.
Experience implementing monitoring based on Golden Signals:
Latency
Traffic
Errors
Saturation
Ability to develop symptom-based, user-impact-focused alerting.
Linux & Networking
Strong knowledge of Linux system administration.
Strong understanding of:
DNS
TCP/IP
Load Balancing
SSL/TLS
Networking fundamentals
Experience supporting highly available production environments.
Incident & Reliability Engineering
Ability to rapidly diagnose and resolve high-severity production incidents.
Experience driving MTTR reduction.
Strong debugging and analytical problem-solving skills.
Ability to identify recurring issues and implement permanent corrective/preventive solutions.
Good-to-Have Skills
Experience working across Azure + AWS + GCP in a multi-cloud environment.
Knowledge of Go (Golang).
Experience with OpenSearch / ELK Stack.
Experience supporting AI/ML workloads in production.
Exposure to Azure AI Services and Azure AI Foundry.
Experience supporting RAG (Retrieval-Augmented Generation) workloads.
Experience designing infrastructure for AI/ML platforms.
Experience building enterprise-wide OpenTelemetry observability frameworks.
Strong understanding of distributed systems architecture.
Exposure to advanced cloud-native architectures and reliability patterns.
Experience with security, compliance, vulnerability remediation, secrets management, and network segmentation.
Key ResponsibilitiesProduction & Incident Management
Participate in the 24/7 on-call rotation.
Diagnose, mitigate, and resolve production incidents.
Lead RCA and post-incident reviews.
Implement corrective and preventive actions.
Continuously improve MTTR and production stability.
Reliability Engineering
Define and improve SLIs, SLOs, SLAs, and Error Budgets.
Identify and eliminate operational toil.
Conduct reliability and capacity reviews.
Improve redundancy, failover, disaster recovery, and system resilience.
Cloud & Infrastructure
Manage and optimize cloud infrastructure across Azure, AWS, and/or GCP.
Manage Kubernetes clusters and containerized applications.
Implement and maintain Infrastructure as Code using Terraform.
Support CI/CD and Git-based development workflows.
Observability & Performance
Build and improve monitoring, logging, metrics, and tracing.
Implement OpenTelemetry and distributed tracing.
Establish Golden Signals-based monitoring and alerting.
Identify and resolve infrastructure and application performance bottlenecks.
Security
Implement cloud security best practices around IAM, network segmentation, and secrets management.
Support vulnerability remediation and compliance initiatives.
Collaborate with Development, Security, and Infrastructure teams.
Ideal Candidate
We Are Looking For Someone With
Strong SRE mindset and production ownership.
Excellent troubleshooting and incident-management skills.
Hands-on expertise in Cloud + Kubernetes + Terraform + Observability.
Strong understanding of OpenTelemetry and Golden Signals.
Experience working in highly available, production-critical environments.
Ability to remain calm and make effective decisions during critical incidents.
Strong communication and cross-functional collaboration skills.
Passion for automation, scalability, reliability, and continuous improvement.
Important Hiring Criteria
Must Be
7+ years relevant experience
Immediate joiner
Willing to work from office in Viman Nagar, Pune
Comfortable with 3:00 PM – 12:00 AM shift
Comfortable with 24/7 on-call rotation
Strong hands-on SRE/DevOps experience
Strong Cloud + Kubernetes + Observability experience
Strong production incident management experience
Good To Have
Multi-cloud: Azure + AWS + GCP
OpenTelemetry
AI/ML or RAG production workloads
Azure AI / AI Foundry
Go
OpenSearch / ELK
Distributed systems
Skills: sre & production operations,gcp,sre,github,mttr,rca,golang,production engineering,sli,cloud infrastructure,golden signals,capacity planning,gitlab,ms azure,iac,devops,toil reduction,opentelemetry,kubernetes & containerization,error budgets,linux system,aws,incident management,terraform,python,sla,24/7 production support,troubleshooting,slo

Work arrangement
No

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App