Live opening · Posted 6 hours ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Domain: Cloud | Kubernetes | Data Platforms | Platform Engineering
Cogrion is building an Autonomous Data & AI Infrastructure Platform for modern enterprises.
We’re looking for a Senior DevOps Engineer who can design, automate, secure, and operate large-scale cloud-native platforms across Kubernetes, data infrastructure, and AI workloads.
What You’ll Work On
Design and operate production-grade Kubernetes platforms
Build and manage infrastructure using Terraform / OpenTofu and Helm
Implement GitOps and CI/CD using ArgoCD, GitHub Actions, GitLab CI, or similar
Manage workloads across AWS, Azure, GCP, and Alibaba Cloud
Design autoscaling, capacity management, and cost optimization strategies
Operate platforms running Spark, Trino, Airflow, Jupyter, MLflow, Kafka, and related services
Build observability using Prometheus, Grafana, Loki, OpenTelemetry, and alerting systems
Improve platform reliability, availability, security, and disaster recovery
Implement IAM, workload identity, secrets management, network security, and RBAC
Troubleshoot complex Kubernetes, networking, storage, and distributed-system issues
Automate platform deployment, upgrades, patching, backup, and recovery
What We’re Looking For
7–10 years of DevOps / SRE / Platform Engineering experience
Strong Kubernetes and Docker expertise
Strong experience with AWS and at least one additional cloud platform
Terraform / OpenTofu, Helm, and GitOps
CI/CD design and automation
Linux, networking, DNS, ingress, load balancing, and cloud networking
IAM, security, secrets, certificates, and workload identity
Strong scripting skills in Python, Bash, or similar
Production troubleshooting and incident-management experience
Strong understanding of scalability, reliability, and infrastructure cost optimization
Bonus
Experience with:
EKS / AKS / GKE / Alibaba ACK
Karpenter and Kubernetes autoscaling
Spark and distributed data workloads on Kubernetes
Trino, Airflow, JupyterHub, MLflow, Kafka
Keycloak, OIDC, IRSA / workload identity
Prometheus, Grafana, Loki, OpenTelemetry
ArgoCD
Multi-tenant SaaS or enterprise platforms
SOC 2 / ISO 27001 / cloud security practices
What Matters to Us
We’re looking for someone who can go beyond infrastructure operations and think like a platform engineer and product owner.
You should be comfortable taking ownership from:
Architecture → Automation → Security → Deployment → Observability → Production Reliability
At Cogrion, you’ll work at the intersection of:
Cloud Infrastructure × Kubernetes × Data Platforms × AI Infrastructure
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.