Live opening · Posted 12 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Principal Cloud Platform Architect | Remote – US
Overview
We’re supporting a rapidly scaling AI infrastructure company building a next-generation cloud platform for large-scale AI training and inference.
They’re entering a major greenfield build phase, creating a platform capable of deploying and operating GPU infrastructure across multiple data centre locations. This is a rare opportunity to define how the entire platform fits together from bare metal and networking through Kubernetes, storage and observability.
They’re looking for a Principal Cloud Platform Architect to own the end-to-end technical architecture and ensure multiple engineering workstreams converge into one scalable platform.
The Opportunity
Own the architecture of a greenfield AI cloud platform from POC through production.
Work across bare metal, Kubernetes, networking, storage and GPU infrastructure.
Solve complex integration challenges across multiple specialist engineering teams.
Influence major build-vs-buy and technology decisions across the platform.
Help create a repeatable architecture for deploying new AI infrastructure sites.
Drive practical adoption of AI coding tools and agentic engineering workflows.
The Role
You’ll sit across several specialist infrastructure teams, connecting their architectural decisions and ensuring the complete system works together.
You won’t replace the individual domain leads. Your focus is the architecture between those domains: identifying integration risks, challenging technical decisions and ensuring the platform can scale into production.
Responsibilities
Own end-to-end architecture and integration across the cloud infrastructure stack.
Define interfaces between bare metal provisioning, Kubernetes, networking, storage and the control plane.
Identify architectural risks before they become production issues.
Lead technical reviews across major platform and infrastructure decisions.
Evaluate infrastructure vendors and guide build-vs-buy decisions using hands-on evidence.
Ensure individual engineering workstreams converge into a coherent production platform.
Help define a repeatable architecture for deploying infrastructure across future sites.
Drive adoption of AI-assisted engineering, automation and agentic workflows.
Skills & Experience
Essential
Experience designing or operating large-scale cloud, GPU, HPC or Kubernetes infrastructure.
Strong architecture experience across multiple infrastructure domains rather than one isolated layer.
Deep technical knowledge in at least three of Kubernetes, networking, bare metal provisioning and GPU/HPC infrastructure.
Experience designing platforms from the ground up in greenfield environments.
Strong understanding of infrastructure integration across multiple engineering teams.
Experience evaluating infrastructure technologies and making build-vs-buy decisions.
Practical experience using AI-assisted development tools as part of engineering workflows.
Nice to Have
Bare metal provisioning experience with technologies such as Redfish, PXE, Metal3 or Ironic.
Kubernetes experience including Cluster API, operators or multi-cluster environments.
Networking knowledge across EVPN/VXLAN, InfiniBand or RoCE.
GPU infrastructure experience including NVIDIA or AMD accelerator environments.
Experience building infrastructure platforms used by external customers.
Interested?
Apply directly or message me for a confidential discussion to learn more.
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.