Live opening · Posted 3 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Responsibilities:
Cloud Architecture and Orchestration: Design and manage multi-cloud environments (AWS/GCP) using Terraform to ensure infrastructure is versioned, reproducible, and scalable.
Kubernetes Mastery: Oversee production-grade Kubernetes clusters, focusing on cluster health, resource optimisation, and seamless application deployment.
AI and Data Infrastructure: Build and maintain the specialised infrastructure required for AI workloads.
Database Management: Maintain a diverse data layer, including Vector search databases (for AI retrieval), PostgreSQL, MySQL, and MongoDB.
Observability and Reliability: Implement deep-stack monitoring and alerting using Datadog and Prometheus to ensure proactive issue detection and resolution.
Automation-First Mindset: Maintain and evolve an active codebase in Python, Go, or Bash to automate repetitive tasks. You will also integrate LLMs into your workflow to accelerate scripting, documentation, and operational efficiency.
Networking and Traffic: Manage complex cloud networking topologies, including VPCs, Load Balancing, Service Meshes, and Caching layers (e. g., Redis) to minimise latency.
Incident Response: Lead the debugging of complex, distributed systems issues, performing root cause analysis to prevent recurrence, as part of an on-call rotation.
Requirements:
Experience: 7+ years of experience in Infrastructure, DevOps, or Site Reliability Engineering, with at least 5 years focused on cloud-native environments.
Linux Systems: Solid understanding of Linux administration and performance tuning. You are comfortable navigating the command line to diagnose system-level issues.
Infrastructure as Code: Expert-level experience with Terraform (or OpenTofu) and managing state at scale.
Containerization: Proven track record of managing Kubernetes in a production environment (EKS, GKE, or self-managed).
Software Engineering: Strong proficiency in Python or Go. You treat infrastructure code with the same rigour as application code (testing, PR reviews, CI/CD).
Data Systems: Hands-on experience managing relational and non-relational databases.
PaaS & Serverless: Experience leveraging and securing PaaS offerings to speed up development cycles.
AI-Augmented Engineering: Practical experience using LLMs (GitHub Copilot, ChatGPT, Claude) to increase your personal and team productivity.
Experience
7-11 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.