Live opening · Posted 22 hours ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Senior Site Reliability Engineer (SRE)
Experience: 4+ Years
Role Type: Full-time
Deployment: Client Location (Bengaluru) / Client Production Environment
Work Schedule: Shift-based, primarily aligned to US time zones
Who We Are
CodeXray is building a modern observability and reliability platform for complex, high-volume production environments.
We work across observability, cloud infrastructure, distributed systems, automation, and AI-assisted SRE, helping enterprises improve production visibility, troubleshooting, incident response, and reliability.
For this role, you will be deployed as part of the CodeXray SRE team at a client environment, working closely with client engineering and operations teams on business-critical production systems.
The Role
We are looking for a hands-on Senior SRE with strong production ownership, troubleshooting depth, and the ability to operate confidently in high-volume AWS environments.
This is not a monitoring or ticket-closure role.
You will be expected to lead critical incidents, run war rooms, troubleshoot across multiple technology layers, guide L1 SREs, and drive issues through resolution, RCA, and preventive action.
What You Will Own
Lead critical production incidents and war rooms
Troubleshoot issues across AWS, Kubernetes, Linux, applications, networks, databases, and dependent services
Own incidents from detection through mitigation, resolution, RCA, and follow-up actions
Lead and guide a team of L1 SREs / Operations Engineers
Work closely with client application, infrastructure, cloud, database, and engineering teams
Correlate metrics, logs, traces, and infrastructure signals to identify root causes
Improve monitoring, alerting, runbooks, SOPs, escalation paths, and operational readiness
Ensure clear communication and structured handovers across shifts
Identify recurring issues and drive corrective and preventive actions
What We Are Looking For
4+ years of experience in SRE, Production Engineering, DevOps, Cloud Operations, or similar roles
Strong hands-on troubleshooting experience in AWS production environments
Experience supporting high-volume, distributed, business-critical systems
Strong knowledge of Linux, networking, containers, and Kubernetes
Good understanding of DNS, TCP/IP, HTTP/HTTPS, load balancers, proxies, and connectivity troubleshooting
Experience with observability tools such as CloudWatch, Prometheus, Grafana, OpenTelemetry, ELK/OpenSearch, or similar
Strong understanding of incident management, escalation, RCA, and problem management
Ability to troubleshoot across multiple layers rather than depend only on dashboards or runbooks
Strong communication skills and confidence working directly with client stakeholders
What Matters to Us
We value engineers who are:
High on intent, ownership, and accountability
Highly sensitive to production impact
Comfortable taking charge during critical incidents
Able to establish a troubleshooting path when the root cause is unclear
Evidence-driven and methodical under pressure
Strong enough technically to guide L1 engineers while remaining hands-on
Comfortable working in a client-facing environment
What You Will Gain
This role provides exposure to a large-scale enterprise production environment where reliability and operational discipline genuinely matter.
You will get the opportunity to:
Work hands-on with high-volume AWS production systems
Build deep expertise in SRE, cloud troubleshooting, Kubernetes, and observability
Lead real production incidents and complex war rooms
Work closely with senior client engineering and operations teams
Gain exposure to enterprise-scale operational processes and reliability practices
Work with CodeXray’s observability and AI-assisted RCA capabilities
Grow toward broader SRE leadership, production engineering, or reliability architecture roles
Shift & Client Location Requirement
This is a client-location deployment and candidates must be comfortable working from the client office in Bengaluru as required.
The role is shift-based and will have significant alignment to US working hours.
Candidates must be comfortable with evening/night shifts in India, rotational schedules, and critical incident support as required.
Willingness to work from the client location and primarily in US-aligned shifts is mandatory.
Good to Have
Experience with AWS services such as EC2, EKS, RDS, ALB/NLB, Route 53, IAM, S3, and CloudWatch, along with exposure to Kafka, Redis, PostgreSQL/MySQL, Terraform, CI/CD, or other distributed systems.
How We Think About SRE
Understand impact → establish the troubleshooting path → lead the war room → correlate signals → drive resolution → complete RCA → prevent recurrence.
If you enjoy being close to production and taking ownership of difficult operational problems, we would like to speak with you.
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.