Live opening · Posted 7 hours ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Job Description
About the role:
Production reliability is a board-level priority, and we are placing it in one hands-on leader. We are hiring a Vice President of Engineering to found and lead our Site Reliability Engineering (SRE) platform organization and own the mission end to end: raise observability maturity and materially reduce production incidents across the engineering estate. You will do it by making agentic engineering the core method of the function — using AI agents to run the platform itself, and putting AI agents in the hands of product engineers so they can meet the company's observability requirements and SRE goals with far less manual effort.
This is a builder's role with executive visibility. You will stand up a central platform team that treats reliability and observability as a product, deliver an aggressive multi-phase roadmap, and change how thousands of engineers instrument, operate, and take ownership of what they ship — in a large, complex, tech-debt-heavy environment where the winning strategy is paved roads over mandates. Central to that is a two-part AI mandate: making this function fully agent-native, and delivering AI agents to product engineers so they can fulfill their observability requirements and the company's SRE goals with far less manual effort.
You will report directly to the SVP of Platform Engineering and partner with leaders across engineering, product, and the executive team. The mission has a named executive sponsor and committed funding; your job is to turn that mandate into outcomes.
The AI mandate: two transformations you will lead:
Agentic engineering is not a side initiative in this role — it is central to how the function operates and to the value it delivers to the enterprise. You will be accountable for two connected transformations:
1. Transform this function to be fully agent-native
Re-found the platform organization around agentic engineering. Agents — not manual toil — should carry the load of instrumenting legacy code, investigating incidents, generating configuration, assisting on-call, and, over time, executing guarded remediation. You will build the agent control plane, guardrails, and evaluation harnesses that make this safe, and turn the function into the company's proof point for what disciplined, agent-first engineering looks like at scale.
2. Put AI agents in the hands of product engineers for observability and SRE outcomes
Deliver AI agents to product engineers as part of the paved road so they can meet the company's observability requirements and SRE goals with far less manual effort — agents that instrument their services, generate SLOs, dashboards, and alerts, investigate incidents, and assist on-call. These agents, and the reliability, evaluation, guardrails, and cost governance behind them, are how product teams hit reliability targets at scale. This role owns the SRE platform: the agents it provides serve observability and reliability, not product-feature development.
What you'll own:
The SRE platform organization — a central platform team plus embedded/partner SREs, an enablement function, and a cross-cutting reliability champions network.
The observability & reliability platform — telemetry pipeline (OpenTelemetry), metrics/logs/traces backends, dashboards and alerting, the SLO and error-budget system, incident management, and RCA/correlation.
The paved road — shared instrumentation SDKs, templates, and dashboards/alerts/SLOs-as-code that make golden-signal observability near-automatic.
The agentic engineering stack — internal agents for instrumentation, investigation/RCA, config generation, on-call, and guarded remediation — plus the agent control plane, guardrails, and evaluation that keep them safe.
The agentic paved road for product engineers — AI agents delivered to product teams to instrument services, generate SLOs/dashboards/alerts, investigate incidents, and assist on-call — so they meet observability requirements and SRE goals with minimal manual effort, backed by guardrails and evaluation.
Reliability & AI governance — error-budget policy, blameless incident and postmortem practice, AI safety and human-in-the-loop controls, and the metrics reported to senior leadership.
The roadmap, budget, and vendor strategy — an ~18-month phased plan, the operating budget, tooling selection, and cost governance for both telemetry and AI workloads.
Key responsibilities:
Set the strategy and vision. — Own the enterprise strategy for reliability, observability, and agentic engineering. Define what “reliability as a product” and “agent-native engineering” mean here, and keep both tied to business outcomes.
Build and lead the organization. — Recruit, structure, and grow a high-performing platform organization spanning SRE, platform, and AI engineering — hiring and developing senior, staff, and principal talent and the managers who lead them.
Deliver the roadmap. — Execute the phased plan — mobilize and instrument, build foundations, scale and standardize the paved road with SLO coverage and error-budget policy, then bring proactive and agentic capabilities to production — on aggressive, overlapping timelines.
Make the function agent-native. — Drive adoption of internal agents across the platform's own work, with least-privilege access, blast-radius limits, human-in-the-loop controls, and evaluation before any increase in autonomy.
Put agents in product engineers' hands. — Deliver AI agents through the paved road that let product engineers meet observability requirements and SRE goals — instrumenting services, standing up SLOs, and resolving incidents — with far less manual effort, backed by guardrails and evaluation.
Make reliability measurable. — Stand up SLOs and error budgets, modern incident management, and blameless postmortems; establish error-budget policy in partnership with leadership.
Drive adoption across the estate. — Win teams over with paved roads and lighthouse wins, not mandates — through reliability reviews, enablement, office hours, and published before/after results.
Own the tooling and its economics. — Select and evolve an industry-leading, OpenTelemetry-native and AI-native tool stack, with cost and cardinality governance built in from the start for both telemetry and inference.
Operate as an executive partner. — Manage stakeholders across engineering and product, report progress and reliability/AI metrics to the CTO and executive team, and steward budget and headcount.
What success looks like:
First 90 days — organization mobilized and sponsor alignment confirmed; reliability baseline established; 2–4 lighthouse services selected; instrumentation underway with the first golden-signal dashboards and SLOs live; the first internal agents piloted.
By 12 months — the paved road is self-service and adopted by tier-1 teams; SLOs and error-budget policy are in effect; on-call load and alert noise are measurably down; the function is operating agent-first; and product engineers are using the SRE agents to instrument and meet SLOs with far less manual effort.
By 18 months — proactive and guarded agentic capabilities are in production; the program hits its targets — a significant reduction in Sev1/Sev2 incidents, mean-time-to-resolution cut substantially, full SLO coverage on tier-1 services, and a healthier on-call — and the company's SRE goals are being met at scale through agents that product engineers rely on.
Required Qualifications:
15+ years in software engineering, including 7+ years in senior engineering leadership leading SRE, platform, infrastructure, or AI organizations at scale — including managing managers.
A proven track record of improving reliability and reducing production incidents across a large, complex, multi-team estate (thousands of engineers and/or services).
Demonstrated experience building and operating production AI and/or agentic systems at scale — with a working command of LLMs, agent frameworks and orchestration, retrieval, evaluation (evals), guardrails, and AI/agent observability.
Experience establishing AI governance and safety for production agents — least-privilege access, human-in-the-loop controls, blast-radius limits, and evaluation-gated autonomy.
Deep expertise in observability (OpenTelemetry; metrics, logs, traces), SLOs and error budgets, incident management, and cloud-native infrastructure (Kubernetes, infrastructure-as-code).
Experience running an internal platform as a product — paved roads, developer experience, and adoption measured by usage, not decree — and enabling other teams to build on it.
A demonstrated ability to drive org-wide change through influence in a large, matrixed organization rather than through mandate.
Excellent executive communication — able to translate reliability and AI strategy into business terms and present to C-level stakeholders.
Bachelor's degree in Computer Science or a related field, or equivalent practical experience.
Preferred Qualifications:
A track record of building internal AI/agent developer tooling that other engineering teams adopted at scale.
Experience in financial services or another regulated, high-availability, high-compliance envir
More openings worth a look
Recently tracked roles with full details and direct application links.