Live opening · Posted 6 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Job Summary
We are looking for a skilled Cloud / DevOps Infrastructure Engineer to design, implement, automate, and maintain scalable and secure cloud infrastructure across AWS, Azure, and/or GCP.
The ideal candidate will have strong hands-on experience with Kubernetes, Docker, Terraform, CI/CD, Linux, cloud infrastructure, networking, security, monitoring, and troubleshooting. Experience supporting AI/ML workloads, GPU infrastructure, or high-performance computing environments will be an added advantage.
The candidate will work closely with development, engineering, security, and operations teams to build reliable infrastructure and automate deployment and operational processes.
Key Responsibilities
Design, deploy, and manage cloud infrastructure across AWS, Azure, and/or GCP.
Build and manage containerized applications using Docker and Kubernetes.
Develop and maintain Infrastructure as Code (IaC) using Terraform.
Design and maintain CI/CD pipelines for automated application deployment.
Manage source control and development workflows using Git.
Provision, configure, and maintain Linux-based servers and environments.
Implement cloud infrastructure automation to improve scalability, reliability, and operational efficiency.
Configure and maintain cloud networking components including VPC/VNet, subnets, routing, load balancers, DNS, firewalls, and security groups.
Apply cloud security best practices including IAM, access control, secrets management, encryption, and network security.
Monitor infrastructure, applications, and services using appropriate monitoring and logging tools.
Troubleshoot infrastructure, networking, deployment, performance, and production issues.
Support high-availability, scalability, backup, disaster recovery, and business continuity requirements.
Collaborate with application developers and engineering teams to improve deployment and infrastructure processes.
Identify opportunities to automate repetitive operational tasks.
Participate in incident management, root-cause analysis, and resolution of production issues.
Maintain infrastructure documentation, deployment procedures, and operational runbooks.
Implement best practices for cloud cost optimization, resource utilization, and infrastructure performance.
Required Technical Skills
Cloud Platforms
Hands-on experience with one or more:
AWS
Microsoft Azure
Google Cloud Platform (GCP)
Containers & Orchestration
Strong experience with Docker.
Hands-on experience with Kubernetes.
Knowledge of Kubernetes deployments, services, ingress, config maps, secrets, namespaces, and scaling.
Infrastructure as Code
Strong experience with Terraform.
Experience creating and managing reusable infrastructure modules.
Understanding of Infrastructure as Code principles and automated provisioning.
CI/CD & Version Control
Experience building and maintaining CI/CD pipelines.
Strong knowledge of Git and Git-based development workflows.
Experience with tools such as Jenkins, GitHub Actions, GitLab CI/CD, Azure DevOps, or similar platforms.
Linux
Strong Linux administration and troubleshooting skills.
Experience with shell scripting and system-level troubleshooting.
Understanding of processes, services, permissions, networking, storage, and system performance.
Networking & Security
Strong understanding of:
TCP/IP
DNS
HTTP/HTTPS
Load Balancing
Firewalls
VPN
VPC/VNet
Subnets
Routing
Security Groups / Network ACLs
IAM and access management
Monitoring & Troubleshooting
Experience with infrastructure and application monitoring.
Knowledge of logging, alerting, metrics, and performance monitoring.
Ability to troubleshoot production issues across cloud, networking, Kubernetes, Linux, and application infrastructure.
Preferred / Good-to-Have Skills
Experience supporting AI/ML infrastructure.
Exposure to GPU infrastructure and GPU-enabled workloads.
Experience with NVIDIA GPUs, CUDA, or GPU scheduling is a plus.
Experience deploying and managing ML/AI workloads on Kubernetes.
Knowledge of MLOps or AI platform infrastructure.
Experience with Helm and Kubernetes package management.
Experience with Prometheus, Grafana, ELK/EFK, Datadog, CloudWatch, Azure Monitor, or similar tools.
Experience with Python, Bash, or other scripting languages.
Knowledge of cloud cost optimization and FinOps practices.
Qualifications
Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
4–8 years of experience in Cloud, DevOps, Infrastructure, SRE, or related engineering roles.
Strong problem-solving and troubleshooting skills.
Ability to work independently as well as collaboratively with cross-functional teams.
Strong communication and documentation skills.
Ideal Candidate Profile
The ideal candidate is a hands-on infrastructure engineer who can build, automate, secure, monitor, and troubleshoot modern cloud environments. The candidate should be comfortable working across cloud platforms, Kubernetes, Terraform, CI/CD, Linux, networking, and security.
Experience with AI/ML platforms or GPU-based infrastructure will be highly valuable for supporting next-generation AI workloads.
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.