Live opening · Posted 7 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
The Executive Director, Site Reliability Engineering owns reliability for a major product domain of the global network estate, as a senior member of the Network Services SRE leadership team. This leader drives the definition and adoption of SLOs for the domain, owns and prioritises its reliability backlog, runs its reliability and post-incident reviews, and directly leads the SREs embedded in the domain alongside domain engineers aligned to the SRE practice through a dotted line. Working with the SRE leads of the other domains, they define and drive SRE culture, standards, and ways of working across the whole of Network Services.
In parallel, the Executive Director serves as the Argentina (Buenos Aires) Regional Lead for Infrastructure Platform Foundational Services (Network, Storage, DataCentre, Data Protection and Recovery). Accountable for site strategy, talent development, governance, executive presence, and operational alignment across the region. The role establishes Argentina as a strategic engineering hub by building and sustaining high-performing engineering teams aligned to global priorities and outcomes.
This leader partners closely with network engineering, Network Rapid Response (NRR), Product Engineering, and the automation platform and data engineering teams to ensure strong operational alignment, effective escalation, end-to-end service ownership, and continuous reliability improvement.
Job responsibilities
Functional Responsibilities - Site Reliability Engineering
Own the reliability outcomes for a product domain of the network estate, and build deep understanding of its architecture, failure modes, and risk profile
Drive the definition, instrumentation, and adoption of SLIs and SLOs across the domain's services, working through the domain's engineering teams rather than defining them in isolation, and use error-budget consumption as a live prioritisation input
Own the domain's reliability backlog: identify the engineering work that materially improves availability, detection, and recovery, keep it prioritised, and hold it visible to domain leadership
Run the domain's reliability reviews and lead its post-incident practice, ensuring blameless review, credible systemic analysis, and that resulting engineering fixes are tracked through to landing
Partner with the network engineering leads in the domain to prioritise reliability work against product engineering and service delivery demand, and make the trade-offs explicit to stakeholders
Directly manage the SREs assigned to the domain, and lead domain engineers aligned to the SRE practice on a dotted line, holding both populations to the same standards, tooling, and career framework
Work with the SRE leads of the other domains to define, evolve, and drive SRE culture, engineering standards, common tooling, the hiring bar, and the career framework across Network Services
Share patterns, tooling, and lessons learned across domains so reliability improvements compound rather than being rebuilt in isolation
Translate the domain's reliability needs into requirements for the automation platform and data engineering teams, and drive adoption of the resulting capability in the domain
Drive reduction of toil in the domain: treat recurring manual work and repeat failure as engineering defects, with measurable reduction targets
Drive adoption of AI and agentic capability across the reliability lifecycle (triage, diagnosis, root-cause analysis, remediation, AI-assisted development) with clear validation standards, so speed never compromises correctness, security, or risk
Hold a code-first engineering bar: production-quality code, testing, code review, and CI/CD, so reliability work ships as software rather than scripts
Deliver measurable outcomes for the domain: improved availability, reduced detection and recovery times, fewer repeat incidents, reduced manual touch, and improved change success rate
Regional Responsibilities - Argentina IP Foundational Services Site Lead
Represent the Argentina hub in global Infrastructure Platforms Foundational Services forums
Drive site culture, retention, and consistent engineering standards across functional silos
Ensure effective cross-functional collaboration with security, infrastructure, application teams, service management, and business partners to deliver integrated outcomes
Manage regional resource planning and budget inputs: capacity forecasting, skills coverage, on-call sustainability, and investment recommendations tied to measurable service improvements
Hire, develop, and retain engineering talent across the Argentina footprint
Ensure governance, risk, and audit compliance for in-country operations
Leadership Expectations
Strong ownership of reliability outcomes, with the technical depth to lead engineers rather than only manage them
Ability to build and scale engineering teams across functional and matrixed reporting lines, including dotted-line reports
Influences peers and partner teams without direct authority, and drives change from within a leadership team rather than from the top of one
Drives the transformation from an operations model to an engineering model
Strong executive communication skills, including calm, credible communication during high-severity events
Maintains compliance and control discipline
Credible senior technology presence in-region, able to represent JPMC externally with regulators, universities, and partners
Required Qualifications, Capabilities, and Skills
10+ years in infrastructure, production, or reliability engineering, including leadership roles running systems at scale
Demonstrated ownership of SLI/SLO/error-budget practice, incident and post-incident leadership, and measurable toil reduction in a pr
More openings worth a look
Recently tracked roles with full details and direct application links.