Live opening · Posted 5 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
About the Role
We are looking for a Consultant / Advisor, Inference Serving and Solution Architecture who has already designed, built, and operated large-scale production LLM inference infrastructure and can bring that experience to what we are building at Firebird AI.
This is a senior technical architecture role for someone who has operated these systems at scale, understands the trade-offs behind different serving architectures, and can provide clear technical direction based on real production experience—including what worked, what failed, what it cost, and what they would do differently.
You will help define the architecture, technology choices, operating model, and delivery approach for our inference platform during its critical early phases. You will work closely with ML Engineering, Platform, Infrastructure, and product teams, establishing reference architectures and engineering practices that can ultimately be owned and evolved by our in-house teams.
This role goes significantly deeper into inference architecture, serving infrastructure, capacity planning, and technical decision-making than a traditional ML Engineering position.
What You'll Do
LLM Inference Architecture
Define the architecture and technical direction for large-scale production LLM inference services.
Evaluate and make architectural decisions across serving technologies including vLLM, TensorRT-LLM, SGLang, llm-d, NVIDIA Dynamo, and related inference infrastructure.
Design gateway and routing architectures, runtime selection strategies, workload placement, and serving patterns appropriate for different models and use cases.
Define approaches to quantization, GPU utilization, fleet sizing, capacity planning, throughput, latency, and cost optimization.
Develop models for capacity requirements and cost per token, helping guide infrastructure and commercial decisions.
Architecture & Technical Decision-Making
Develop reference architectures, Architecture Decision Records (ADRs), technical standards, and design principles.
Evaluate architectural alternatives and clearly articulate trade-offs, rejected approaches, risks, and long-term consequences.
Provide senior technical guidance during critical design and implementation decisions.
Review designs and challenge assumptions across ML, Platform, Infrastructure, and product engineering.
Modernization & Migration
Define strategies for evolving and modernizing the inference stack as technologies, models, and infrastructure requirements change.
Lead or advise on migrations of live production workloads to new serving stacks and architectures while minimizing operational risk.
Establish approaches for compatibility, rollout, validation, rollback, and production readiness.
Delivery & Operating Model
Help define the ML delivery lifecycle, including development, testing, staging, production, and promotion strategies.
Establish release and production-readiness criteria for inference services.
Define clear ownership boundaries between model, ML engineering, platform, infrastructure, and product teams.
Help establish engineering and operational practices that allow the platform to scale sustainably.
Transfer architectural knowledge, standards, and decision-making frameworks to internal engineering and architecture teams.
What You'll Bring
8+ years of engineering experience, including at least 3 years operating at Staff, Principal, Architect, or equivalent level with meaningful technical design authority.
Significant hands-on experience designing, building, and operating LLM inference services in production at scale.
Direct ownership of production inference systems and a strong understanding of their reliability, performance, operational, and cost characteristics.
Deep technical knowledge of modern LLM serving stacks such as vLLM, TensorRT-LLM, SGLang, llm-d, NVIDIA Dynamo, and related technologies.
Strong understanding of inference routing, runtime selection, capacity and throughput planning, GPU fleet sizing, quantization strategies, and cost-per-token modeling.
Strong solution architecture experience, including reference architectures, ADRs, architecture reviews, and technology evaluation.
Experience modernizing or migrating live production systems to new architectures or technology stacks without disrupting service.
Experience defining engineering delivery processes, environment and promotion strategies, release criteria, and production ownership models.
Ability to make technically rigorous decisions while clearly communicating trade-offs to engineering, product, and business stakeholders.
Nice to Have
Experience designing multi-node and multi-region inference architectures, disaster recovery strategies, and capacity commitment models.
Experience operating inference workloads across on-premises GPU data centers and cloud infrastructure.
Experience supporting pre-sales activities, including solution design, technical estimation, proposals, and Statements of Work (SoWs).
Production architecture experience in regulated industries and an understanding of how regulatory and audit requirements affect AI serving infrastructure.
Familiarity with SOC 2, ISO 27001, GDPR, and the EU AI Act.
Experience mentoring senior engineers and architects and establishing effective architecture and design-review practices.
Public technical contributions, conference participation, publications, open-source work, or another demonstrated industry track record.
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.