Live opening · Posted 6 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Detailed Job Description
Key responsibilities
Own reliability outcomes for mission-critical network services (availability, performance, recoverability) and drive measurable improvements.
Lead major incident response: establish response rigor, coordinate technical mitigation, coach responders, and ensure high-quality post-incident learning.
Own problem management end-to-end: author/approve RCAs, identify systemic root causes, and drive durable remediation and prevention.
Architect automation frameworks using Python, Shell, and Ansible (e.g., self-healing patterns where appropriate, automated remediation with guardrails, validation pipelines, standardized tooling).
Provide deep technical leadership across:
SD-WAN, SDA, and software-defined networking (SND)
Routing and Switching
Firewalls, Load Balancers, Proxies
Strong preference for Cisco ACI / Fabrics; VMware NSX as a plus
Partner with developers/platform teams to develop observability insights: define operational signals, reduce alert noise, improve actionability, and align monitoring to service health.
Embed SRE engineering practices into design and delivery:
Translate NFRs into controls, standards, and measurable targets
Apply FMEA-style risk analysis to designs/changes to reduce failure impact
Drive operational readiness reviews, resilience testing, and change safety mechanisms
Mentor engineers, set technical standards, influence cross-team roadmaps, and deliver outcomes with minimal supervision.
Required qualifications
Extensive experience operating and engineering large-scale networks with strong troubleshooting depth.
Proven leadership in major incidents, RCA, and delivery of long-term fixes that reduce repeat incidents.
Advanced automation track record using Python, Shell, and Ansible, with demonstrable toil reduction and reliability gains.
Strong applied understanding of SRE concepts, NFRs, and FMEA (or equivalent failure/risk analysis).
Strong ownership mindset, stakeholder management, and ability to drive delivery independently.
Preferred qualifications
Strong expertise with Cisco ACI / Fabrics.
VMware NSX experience (nice-to-have).
Certifications such as CCNP/CCIE (or equivalent), plus other vendor certifications.
Prior experience in financial institutions (advantage).
More openings worth a look
Recently tracked roles with full details and direct application links.