Live opening · Posted 11 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Job Title: Principal Solutions Architect, Post-Sales (EMEA) – AI/ML Infrastructure
Role:
A pioneering Kubernetes-native platform provider, enabling enterprises and neoclouds to orchestrate large-scale GPU infrastructure for AI and ML workloads, is hiring a Principal Solutions Architect to anchor its EMEA post-sales function. This is a senior, customer-facing position sitting at the intersection of distributed systems, GPU infrastructure and MLOps, working with some of the most demanding compute environments in the region. The successful candidate will operate with genuine autonomy, acting as the primary technical authority for strategic accounts across EMEA.
Direct exposure to cutting-edge GPU fabric and large-scale distributed training environments, working hands-on with the infrastructure underpinning frontier AI workloads.
Join at a moment of explosive AI compute demand, with genuine investment and scale behind the business, and a seat that feeds directly into the product and engineering roadmap.
A small, senior Solutions Architecture function built on trust rather than micromanagement, with real autonomy, direct access to product and engineering leadership, and a mentorship remit over junior team members.
Responsibilities:
Design end-to-end AI/ML platform architectures spanning inference, training and data pipelines
Develop reference architectures for GPU cluster deployment, LLM serving and multi-tenant ML infrastructure
Advise on GPU fabric topology (NVLink, InfiniBand, RoCEv2) for distributed training workloads
Act as the primary technical advisor and escalation point, leading root cause analysis on complex production issues
Deliver technical workshops, proofs of concept and executive-level presentations to senior stakeholders
Design observability strategies using DCGM, OpenTelemetry, eBPF and GPU metrics pipelines
Partner with customer platform, MLOps and data science stakeholders to translate requirements into architecture
Feed customer insights back into the product and engineering roadmap, and mentor junior Solutions Architects
Skills/Must have:
Experience: 8+ years in infrastructure, platform or solutions engineering, including 3+ years specifically in AI/ML infrastructure or MLOps
Core Tech/Domain: Deep hands-on Kubernetes expertise (cluster lifecycle, workloads, operators, RBAC) and direct NVIDIA GPU infrastructure experience (H100/H200/B200 preferred)
Methodology/Protocols: Distributed training knowledge (NCCL, tensor/pipeline parallelism), LLM inference serving and optimisation (vLLM, NIM, TGI), and GPU networking (GPU Operator, MIG, SR-IOV, IB/RoCEv2)
Soft Skills: Strong customer-facing and advisory communication skills, credible at both engineer and executive level
Nice to haves:
Run:AI or Slurm scheduling experience
PyTorch/TensorFlow familiarity
Public cloud experience (AWS/Azure/GCP) including networking, IAM and managed Kubernetes
CKA, CKAD, AWS Solutions Architect, Azure Solutions Architect or GCP Professional Cloud Architect certification
Multi-tenant GPU isolation knowledge (SR-IOV VFs, DPU offload)
Benefits:
Performance-related bonus
Private healthcare
Pension contribution
Training and certification allowance
Remote-first, flexible working across EMEA
Salary:
£150,000 to £180,000 base plus performance-related bonus
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.