Live opening · Posted 13 days ago

SRE Guild Lead

EZ INFORMATICS SOLUTIONS PVT LTD · Kolkata metropolitan area, West Bengal, India (Hybrid)
Linkedin No
You are 13 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 13 days ago
CompanyEZ INFORMATICS SOLUTIONS PVT LTD
LocationKolkata metropolitan area, West Bengal, India (Hybrid)
Work modeNo
SourceLinkedin
Listed13 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
10 min from Linkedin publishing this role to us finding it
12 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
71,832 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

Scope of work
We are looking for an experienced SRE Guild Lead to own the practice, standards, and craft of Site Reliability Engineering across our Factory pods. As guild lead, the person will be responsible for the
technical and cultural anchor for reliability at TCG Digital – setting the bar for how we define, measure, and defend service reliability, and ensuring every pod applies consistent SRE practices regardless of which
product or client engagement they sit under. The role will work closely with the DevOps & Infra guild, Architecture & Design guild, and pod leads to embed observability, automation, and operational excellence into how we build and run systems. To succeed in this role, the person should have hands-on production engineering depth combined with the ability to mentor and set standards across a distributed group of engineers who do not report to the role directly. This is a guild leadership role, not a pure peoplemanagement role: influence, technical credibility, documentation, and cross pod coordination are your primary tools. The person will also be the escalation point for major incidents and a key voice in capacity planning for reliability roles across the delivery organization.
Essential Skills :
• Excellent communication and influencing skills –able to set and land standards across pods without formal reporting authority.
• Deep hands-on experience with observability stacks (e.g., Prometheus, Grafana, ELK/OpenSearch, Datadog) – designing dashboards, alerts, and golden signals.
• Proven experience defining and operationalizing SLIs, SLOs, and error budgets, and using them to drive engineering and release decisions.
• Strong background in incident management – oncall design, escalation paths, severity classification, and blameless postmortems.
• Hands-on expertise with Kubernetes, container orchestration, and cloud-native infrastructure (AWS, Azure, or GCP).
• Infrastructure as Code and automation experience (Terraform, Ansible, Helm, or equivalent) with a strong toil-reduction mindset.
• Proficient in scripting/programming for automation and tooling (Python, Go, or Shell).
• Experience with CI/CD platforms and release engineering (Jenkins, GitLab CI, ArgoCD, or similar).
• Working knowledge of capacity planning, performance tuning, and cost optimization for production systems.
• Familiarity with chaos engineering, disaster recovery, and multiregion resilience practices.
• Ability to mentor engineers across multiple pods, run guild syncs, and maintain shared runbooks, standards, and playbooks.
• Proven ability to work cross-functionally with Architecture, DevOps & Infra, and Delivery leadership
Roles & Responsibilities
• Own the SRE guild charter – define reliability standards, tooling, and
best practices applied consistently across all Dev Factory pods.
• Define and govern SLIs, SLOs, and error budgets for key services in
partnership with pod leads and architects.
• Lead the guild's incident management practice – on-call rotations,
escalation paths, and blameless postmortems; act as senior escalation
point for critical production incidents.
• Drive observability strategy across the organization, ensuring pods
have consistent monitoring, alerting, and logging coverage.
• Champion automation and toil reduction, identifying repetitive
operational work across pods and driving it toward self-service or
automated remediation.
• Partner with the DevOps & Infra guild to align infrastructure,
deployment, and reliability practices, and flag capacity or staffing
gaps to guild and delivery leadership.
• Mentor and upskill SREs and DevOps engineers embedded in pods,
running guild-wide knowledge-sharing sessions and maintaining
shared runbooks.
• Contribute to hiring, onboarding, and career pathing for the SRE
discipline across the delivery organization.
• Collaborate with the Architecture & Design guild on non-functional
requirements – scalability, fault tolerance, and disaster recovery –
during solution design.
• Track and report guild health metrics (incident trends, MTTR, SLO
attainment, automation coverage) to Dev Factory leadership.
• Stay current on SRE tooling and industry practices, evaluating and
introducing new technologies where they improve reliability or
reduce operational cost

Work arrangement
No

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App