Live opening · Posted 18 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Responsibilities:
Design and operate a model serving infrastructure across on-prem and cloud deployments.
Build and maintain CI/CD pipelines for model updates, rollbacks, and evaluation-gated deployments.
Monitor model performance in production latency, accuracy drift, throughput, failure modes and build systems that surface issues before clients do.
Build evaluation infrastructure: harnesses, A/B testing, and model comparison tooling for field and lab use.
Manage containerised model serving in constrained, air-gapped, and edge environments.
Collaborate with data scientists on eval pipelines; own the infrastructure layer underneath.
Create runbooks and operational playbooks that strategic deployment engineers can use in the field.
Own incident response for model-layer failures across all active deployments.
Requirements:
3-5 years in ML engineering or MLOps with at least one production LLM or ML system in continuous operation.
Deep expertise in model serving: vLLM, TGI, Triton Inference Server, or equivalent; experience with quantised model formats (GGUF, AWQ, GPTQ).
Experience fine-tuning and adapting models in constrained, on-premises, or air-gapped environments, including managing data pipelines and compute limitations specific to the environment.
Containerisation experience with Docker, Kubernetes, or lightweight alternatives (K3S, K0S) for constrained and edge environments; familiarity with deploying across heterogeneous hardware and infrastructure configurations.
Monitoring and observability using Prometheus, Grafana, or equivalent; ability to build custom eval dashboards.
Python fluency; familiarity with fine-tuning workflows and model evaluation frameworks.
Hands-on experience with CI/CD tooling for ML pipelines: GitHub Actions, ArgoCD, DVC, or similar.
Good to Have:
You've kept a production ML system running under load and debugged it when it broke.
You don't wait for things to fail; you build systems that tell you when they're about to.
You write documentation that actually gets used by people who aren't you.
Experience
3-5 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.