Live opening · Posted 13 hours ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
We're hiring Senior Members of Technical Staff (MTS) who are strong software engineers and want to work close to production, customers, and the underlying systems powering enterprise AI.
You will be mapped to one of three technical tracks based on your strengths:
AI Platforms: Enterprise AI platforms, backend systems, LLM applications and agents
AI Infra: GPU infrastructure, Kubernetes, distributed systems and platform engineering
Applied Research: AI evaluation, post-training, simulations and data curation
You don't need to choose a track upfront. Your technical depth and demonstrated experience will determine the best fit.
You have approximately 5 years of professional software engineering experience and a strong track record of shipping and operating production systems. You should be able to take ownership of a system or significant workstream end-to-end, from architecture and design through implementation, deployment, operations, debugging, and continuous improvement. We value strong engineering fundamentals, practical problem-solving, and the ability to make sound technical decisions in ambiguous environments.
Requirements:
Strong programming ability in Python, Go, or Rust.
Experience designing, building, and operating production services at scale.
Strong understanding of APIs, databases, distributed systems, and software architecture.
Hands-on experience with Docker and Kubernetes.
Experience with at least one major cloud platform: AWS, Azure, or GCP.
Strong understanding of Git, CI/CD, automated testing, and engineering best practices.
Understanding of reliability concepts such as fault tolerance, retries, idempotency, failure handling, logging, monitoring, and observability.
Ability to debug production issues, analyse system behaviour, and drive incidents through resolution.
Ability to evaluate technical trade-offs and make architecture decisions with long-term maintainability in mind.
Track 1 AI Platforms
Best suited for backend and platform engineers who have hands-on experience building production applications and services around LLMs and enterprise AI
You should have experience with several of:
Strong Python backend engineering
REST/gRPC APIs and production services
PostgreSQL / Redis and schema design
Kafka, NATS, or other messaging systems
Event-driven and distributed architectures
LLM applications and production model APIs
RAG pipelines
Agentic applications and tool calling
Vector databases
LangChain, LangGraph, LlamaIndex, CrewAI, AutoGen, or equivalent frameworks
LLM evaluation, observability, tracing, cost, and latency monitoring
Enterprise AI integrations and platform engineering
You will work on
RAG and multi-agent platforms
Inference services
Evaluation and benchmarking systems
Enterprise integrations
Metering and billing
Governance and access control
Audit, privacy, PII handling, and guardrails
AI observability, tracing, and production monitoring
Customer VPC and on-premises deployments
Track 2 AI Infra
Best suited for systems, infrastructure, and platform engineers interested in building large-scale GPU infrastructure for AI workloads
You should have experience with several of:
Strong Go or Rust
Linux systems and infrastructure engineering
Kubernetes and container platforms
Docker / containerd
GPU infrastructure and fleet management
CUDA
NVML / DCGM
MPS / MIG
Kubernetes controllers and operators
REST and gRPC platform APIs
Kafka / NATS / RabbitMQ
PostgreSQL / MySQL
Prometheus / Grafana / OpenTelemetry
Distributed systems and fault tolerance
Networking, resource isolation, and reliability engineering
You will work on
GPU fleet management and control planes
Schedulers and multi-tenant infrastructure
GPU sharing and isolation
GPU health monitoring and recovery
Kubernetes controllers, exporters, daemons, and platform services
On-premises and cloud GPU infrastructure
Infrastructure reliability and observability at scale
Experience with GPU infrastructure, bare-metal systems, Kubernetes platform engineering, or HPC environments is particularly valuable.
Track 3 Applied Research
Best suited for ML engineers who combine strong software engineering fundamentals with hands-on experimentation and a strong bias toward research that reaches production.
You should have experience with several of:
Strong Python and ML engineering fundamentals
PyTorch or JAX
LLM evaluation and benchmarking
Fine-tuning / post-training
Agent evaluation
RL / RLHF
Embedding and reranker experiments
Prompt optimisation and tool-use training
Simulation environments
Synthetic data generation
Large-scale data curation
Model training or serving infrastructure
Statistical experimentation and rigorous evaluation
You will work on
Offline and online evaluation systems
LLM-as-a-judge and regression frameworks
Agent reasoning, planning, and tool-use improvements
Fine-tuning and post-training
Retrieval quality improvements
Simulation environments for agent testing
Data curation and synthetic data pipelines
Production-ready AI research capabilities
What Success Looks Like
As a Senior MTS, you will:
Own systems and significant technical workstreams end-to-end
Drive projects from ambiguous requirements to production
Make architecture decisions and clearly articulate technical trade-offs
Ship reliable production software used by enterprise customers
Operate and improve systems at scale
Debug complex production issues and drive root-cause resolution
Work closely with senior engineers, founders, and customers
Mentor engineers and raise the technical bar within the team
Contribute across engineering, infrastructure, AI, and customer-facing requirements
Continuously expand your technical and business scope
Who Will Stand Out
The strongest candidates will demonstrate:
Real production ownership across design, deployment, operations, and improvement
Strong coding, architecture, and distributed systems fundamentals
Experience operating production systems at scale
Hands-on Docker, Kubernetes, and cloud experience
Strong expertise in at least one of the three Aion tracks
Ability to explain technical decisions, constraints, and trade-offs
Strong debugging and problem-solving skills
Ability to work independently in ambiguous environments
Experience mentoring engineers and influencing technical direction
Strong communication and collaboration skills
A founder-like bias for action and ownership
Ability to connect technical decisions to business and customer outcomes
Comfort working in a fast-moving, early-stage startup environment
Skills
Agentic AI, Architecture, FastAPI, Generative AI, Golang, LangChain, Python
Experience
5-9 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.