Live opening · Posted 11 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
As a Senior AI Platform Ops Engineer, you will join a multidisciplinary team responsible for building and operating the core AI platform capabilities that power internal and customer-facing AI solutions. This role extends traditional DevOps ownership into AI gateway services, model and agent deployment, MCP hosting, MLOps workflows, platform security, and production-grade observability.
Responsibilities:
Design, build, and operate scalable cloud infrastructure and delivery pipelines for AI platform services across development, UAT, and production environments.
Build and support AI platform CI/CD patterns for APIs, models, agents, MCP servers, and supporting services using container-based deployment standards and rollback automation.
Own the runtime operations of AI gateway capabilities, including request routing, authentication, load balancing, retry policies, circuit breakers, and semantic caching.
Build and maintain model deployment patterns, artefact management workflows, and inference pipelines for enterprise AI and ML use cases.
Support agent and MCP platform operations, including standardised templates, health checks, autoscaling, monitoring, RBAC, and service-to-service authentication.
Partner with data, AI, security, and application teams to enable vector databases, graph databases, feature stores, and RAG pipelines as reusable platform services.
Implement platform guardrails such as PII detection, prompt injection protection, data leakage prevention, and use-case-specific business rules.
Build deep observability for AI services, including latency, throughput, error rates, token usage, cost telemetry, audit logs, and operational dashboards.
Drive production reliability through failover design, disaster recovery practices, incident support, runbooks, and continuous operational improvement.
Support API platform integration patterns using Apigee X for secure exposure, routing, policy enforcement, onboarding, and monitoring of AI-related services.
Requirements:
7+ years of experience in DevOps, platform engineering, cloud infrastructure, or site reliability engineering, with strong production operations ownership.
Strong hands-on experience with cloud platforms, microservices, CI/CD, containers, automation, and infrastructure design.
Strong experience with Docker, Kubernetes, Linux administration, Git workflows, and system administration fundamentals.
Hands-on development experience in Python and at least one additional language such as Node.js or Java.
Experience with MLOps concepts such as model packaging, artefact management, deployment pipelines, inference operations, and rollback strategies.
Experience operating AI or data platforms such as Databricks, MLflow, Airflow, or similar enterprise ML tooling.
Strong understanding of observability, logging, alerting, and performance monitoring for distributed systems and AI workloads.
Experience with secrets management, access control, and auditability in enterprise environments.
Working knowledge of Apigee X or similar enterprise API gateway platforms.
Preferred qualifications:
Experience with AI gateway architecture, model catalogues, agent platforms, or MCP hosting patterns.
Experience with vector databases, graph databases, feature stores, and RAG pipeline operations.
Experience with AKS or other managed Kubernetes services, managed identities, Key Vault, TLS, SSO, and enterprise platform runbooks.
Experience working in Agile delivery models with strong cross-functional communication and a bias for action.
Experience
5-9 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.