Live opening · Posted 1 day ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
About the Role
Office Beacon is looking for a Principal AI Platform Architect – Enterprise GenAI to define the technical vision, architecture, and engineering strategy for an enterprise-grade AI platform.
The platform will support secure and scalable deployment of Small Language Models (SLMs), Large Language Models (LLMs), Vision Language Models (VLMs), Retrieval-Augmented Generation (RAG), and AI agent workflows within enterprise environments.
The Principal AI Platform Architect will own end-to-end architectural direction, from technology selection and platform design through implementation, deployment, optimization, governance, and continuous evolution.
This is a hands-on architecture role requiring deep experience building and operating production AI/ML platforms, GPU infrastructure, model-serving systems, distributed systems, and cloud-native infrastructure.
Key Responsibilities
AI Platform Architecture
Define the overall architecture and technical vision for the Enterprise GenAI platform.
Establish architectural standards for AI applications, model serving, data flows, infrastructure, and platform services.
Evaluate technologies and make architecture decisions based on scalability, reliability, security, performance, and maintainability.
Define technical roadmaps for the continued evolution of the AI platform.
Model Serving & GPU Infrastructure
Lead architecture decisions related to GPU infrastructure, model serving, and model lifecycle management.
Design production-grade training and inference environments for AI workloads.
Architect scalable model-serving infrastructure using technologies such as vLLM, Hugging Face TGI, TensorRT-LLM, or similar platforms.
Optimize GPU utilization, inference performance, and infrastructure efficiency.
GenAI, RAG & AI Agents
Architect production-grade RAG systems involving embeddings, semantic search, retrieval, reranking, and vector databases.
Design AI agent architectures and production workflows.
Define architectures supporting LLMs, SLMs, VLMs, and other generative AI workloads.
Establish approaches for model evaluation, deployment, monitoring, and lifecycle management.
Model Adaptation & MLOps
Define approaches for model fine-tuning and adaptation using techniques such as LoRA, QLoRA, and PEFT.
Establish model development, evaluation, deployment, and monitoring workflows.
Implement or guide CI/CD processes for AI models, applications, and infrastructure.
Use appropriate MLOps technologies and practices to improve reproducibility and operational reliability.
Cloud & Infrastructure
Design and operate cloud-native AI infrastructure on AWS, Microsoft Azure, GCP, or comparable platforms.
Architect Kubernetes-based AI workloads and containerized services.
Define infrastructure-as-code practices using Terraform or similar technologies.
Design distributed systems capable of supporting enterprise AI workloads at scale.
Security & Enterprise Architecture
Define security controls for AI workloads and platform infrastructure.
Architect appropriate tenant isolation and data-separation strategies.
Collaborate with Security and DevOps teams on infrastructure and application security.
Contribute to enterprise AI governance and operational standards.
Technical Leadership
Provide technical direction to AI engineers and platform engineering teams.
Review architectural proposals and technical designs.
Establish engineering standards and best practices across AI platform initiatives.
Collaborate with DevOps, Product, QA, Security, and other engineering teams.
Mentor engineers and help resolve complex architectural and technical challenges.
3. Must-Have Qualifications
10+ years of professional software engineering experience.
5+ years of hands-on experience building and operating production AI/ML or Generative AI platforms.
Demonstrated experience designing and deploying enterprise-grade AI systems.
Strong production experience with multiple enterprise GenAI technologies, including LLMs, SLMs, VLMs, RAG, and AI agents.
Strong proficiency in Python and PyTorch.
Hands-on experience with Hugging Face Transformers.
Production experience developing AI services and APIs using FastAPI or similar frameworks.
Strong understanding of LoRA, QLoRA, PEFT, and model fine-tuning.
Strong understanding of embeddings, semantic search, and vector databases.
Hands-on experience with production model-serving or MLOps technologies such as vLLM, Hugging Face TGI, TensorRT-LLM, Ray, or MLflow.
Strong experience with Kubernetes and Docker.
Strong experience with Terraform or another infrastructure-as-code technology.
Strong experience with at least one major cloud platform: AWS, Microsoft Azure, or Google Cloud Platform.
Experience designing or operating GPU-based cloud infrastructure.
Strong knowledge of software architecture and distributed systems.
Proven ability to make architecture decisions for complex production systems.
Preferred Qualifications
Experience designing AI platforms for multi-tenant enterprise environments.
Experience implementing tenant isolation and data-separation strategies.
Experience with SOC 2 or ISO 27001 requirements in technology environments.
Experience with document intelligence and OCR-based AI workflows.
Experience with GPU optimization and inference performance tuning.
Experience with model quantization techniques.
Experience operating large-scale Kubernetes environments.
Experience establishing AI platform governance, observability, and reliability practices.
Experience leading technical architecture across multiple engineering teams.
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.