Live opening · Posted 6 hours ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
About the RoleAs a Sr Director - AI & Platform Engineering you are accountable for building and operating the enterprise agent platform that enables every team to safely develop, deploy, and scale AI agents across the organization. Rather than creating agents directly, the focus is on providing the foundational capabilities that make agent development reliable and governable at scale: the agent authoring experience, tool and MCP integrations, context and memory architecture, skills and procedural knowledge management, and the data access model. You also own the runtime infrastructure, including the model gateway that governs all model traffic and the orchestration engine that supports durable, long-running, and human-in-the-loop workflows.
Equally important, this role establishes the controls and trust mechanisms that ensure agents operate securely, responsibly, and within defined boundaries. Responsibilities include agent identity and delegated authorization, guardrails and policy enforcement, runtime registry controls, and cost and capacity management. The role also oversees evaluation frameworks, promotion and release gates, observability and tracing, deployment management, and continuous improvement processes. When executed successfully, the platform enables dozens of teams to rapidly ship agents on a standardized, governed path while providing complete transparency into every agent action, including what it did, under whose authority it acted, and what evidence informed its decisions.What You'll Do
Define and drive the enterprise agent platform strategy, including architecture, technology choices, and a secure, scalable paved road that accelerates adoption through exceptional developer experience
Lead delivery of the core agent platform, including the model gateway, orchestration runtime, and context/memory services that enable reliable, governed, and cost-effective agent execution at scale
Build and operate the trust and quality framework, including evaluations, observability, release gates, performance monitoring, and continuous improvement mechanisms for production agents
Own platform security, governance, and operational excellence, including agent identity, authorization, policy enforcement, auditability, incident response, SLOs, and cost management
Partner across Engineering, Security, Data Governance, Google Cloud, and executive leadership to align platform capabilities with business priorities, regulatory requirements, and long-term AI strategy
Who You Are
Deep expertise in Google Cloud AI services, including Vertex AI, Gemini, model deployment, vector search, and evaluation frameworks for enterprise-scale AI applications
Strong data platform experience with BigQuery, Spanner, Dataplex, Knowledge Catalog, and governed semantic layers such as Looker/LookML
Proficiency in cloud-native platform engineering, including Cloud Run, GKE, Terraform, CI/CD pipelines, artifact management, and distributed systems architecture
Experience with enterprise security and governance, including IAM, workload identity, secrets management, encryption, policy enforcement, and data protection controls
Knowledge of event-driven and streaming architectures, leveraging technologies such as Pub/Sub, Dataflow, and workflow orchestration platforms
Expertise in observability and reliability engineering, including logging, tracing, monitoring, OpenTelemetry, SLOs, incident management, and production operations
Hands-on experience building agent and LLM platforms, including agent frameworks, MCP integrations, interoperability standards, and large-scale model orchestration
Deep understanding of context, memory, and evaluation systems, including retrieval architectures, memory management, benchmarking, online/offline evaluations, and drift detection
Strong background in AI safety, guardrails, and inference optimization, including prompt-injection defense, grounding verification, multi-model routing, caching strategies, and cost management
Proven software engineering and platform leadership experience, including Python and TypeScript/Go development, multi-tenant platform design, durable workflows, authorization architecture, and secure agent operations
More openings worth a look
Recently tracked roles with full details and direct application links.