Live opening · Posted 7 hours ago

Site Reliability Engineer (SRE) - Observability

Tata Consultancy Services (TCS) · Bengaluru, Karnataka, India (On-site)
Linkedin No
JobBeeper subscribers received an alert for this role.

At a glance

The key details from the original listing.

Posted 7 hours ago
CompanyTata Consultancy Services (TCS)
LocationBengaluru, Karnataka, India (On-site)
Work modeNo
SkillsPython, AWS, Azure, GCP, Docker, Kubernetes, Elasticsearch
SourceLinkedin
ListedPosted 7 hours ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
2 min from Linkedin publishing this role to us finding it
8 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
69,607 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

Site Reliability Engineer (SRE) - Observability
Experience - 8-10 Years
Location - Chennai / Bangalore / Hyderabad / Delhi / Pune
Notice Period - Immediate Joiners to 30 Days Preferred
Job Description :
We are seeking an experienced SRE - Observability Engineer to drive monitoring, logging, tracing, and reliability engineering initiatives across large-scale cloud-native environments. The ideal candidate will possess strong expertise in observability platforms, incident management, performance monitoring, and automation to enhance system availability, reliability, and operational excellence.
Key Responsibilities :
Design, implement, and manage enterprise observability solutions.
Build and maintain monitoring dashboards, alerts, and SLO/SLI frameworks.
Analyze application, infrastructure, and network performance metrics.
Implement and manage logging and distributed tracing solutions.
Conduct incident troubleshooting, root cause analysis (RCA), and postmortem reviews.
Collaborate with Development, DevOps, Cloud, and Platform teams to improve service reliability.
Automate operational tasks using scripting and infrastructure-as-code practices.
Define and optimize alerting strategies to reduce noise and improve incident response.
Required Skills :
Strong experience in SRE and Observability practices.
Hands-on expertise with Prometheus, Grafana, Datadog, Dynatrace, AppDynamics, or New Relic.
Experience with ELK Stack (Elasticsearch, Logstash, Kibana) or Splunk.
Knowledge of OpenTelemetry, distributed tracing, and observability frameworks.
Experience with Kubernetes, Docker, and Cloud Platforms (AWS/Azure/GCP).
Strong understanding of Linux systems, networking, and system performance tuning.
Scripting experience in Python, Shell, or Bash.
Good understanding of Incident Management, Monitoring, and Reliability Engineering concepts.

Work arrangement
No

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App