Live opening · Posted 7 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Leadership & Strategy
Define and implement SRE best practices across the organization.
Proven expertise in production support, engineering, disaster recovery (DCR), automation, and cloud operations
Mentor and guide a team of SREs, fostering growth
Collaborate with senior stakeholders to align reliability goals with business objectives.
Reliability & Performance
Establish SLIs, SLOs, and SLAs for critical services and ensure adherence.
Drive initiatives to improve system and reduce operational toil.
Excellent in designing systems that detect and remediate issues without manual intervention – Self Healing systems, Runbook automation
Exposure to tools like Gremlin, Chaos Monkey, AWS FIS to simulate outages and improve fault tolerance
Incident Management
Act as the primary point of escalation for critical production issues and lead major incident response, root cause analysis, and postmortems.
Perform detailed post-incident investigations to identify underlying causes. Document findings and share learnings to prevent recurrence.
Implement preventive measures and continuous improvement processes.
Observability
Champion monitoring, logging, and alerting strategies using tools like Prometheus, Grafana, ELK, and AWS CloudWatch.
Build real-time dashboards to visualize system health and reliability metrics.
Configure intelligent alerting based on anomaly detection and thresholds.
Combine metrics, logs, and traces to enable root cause analysis and reduce Mean Time to Resolution (MTTR).
Knowledge of AIOps or ML-based anomaly detection for proactive reliability management.
Collaboration
Work closely with development teams to integrate reliability into application design and deployment
Promote a culture of shared responsibility for uptime and performance across engineering teams.
Qualified with a degree in B.Sc. in Computer Science, MCA in Computer Science, Bachelor of Technology in Engineering, or higher
Hands on technologist with minimum 12 years of experience working in software development with at least 5 years of experience leading an SRE team currently
Deep expertise with various AWS services. Advanced knowledge of monitoring and observability tools.
Proven track record of building secure, mission-critical, high-volume transaction web-based software systems, in regulated environments (finance and insurance industries).
Employment type
Full-time
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.