Live opening · Posted 7 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
This is a fully remote role open to candidates across the APAC region
About Saturn Cloud
Saturn Cloud is building infrastructure that helps data science and engineering teams develop, run, and scale AI, machine learning, and analytics workloads in the cloud.
Our platform brings together development environments, orchestration, elastic compute, GPUs, and production infrastructure so teams can focus on their work without having to manage all of the underlying complexity themselves.
We’re looking for a Site Reliability Engineer based in the APAC region to help us build, operate, and continuously improve the infrastructure behind Saturn Cloud.
What You’ll Do
As a Site Reliability Engineer, you’ll work across infrastructure, platform engineering, and operations to keep Saturn Cloud reliable, scalable, and easy to operate.
You’ll:
Build, operate, and improve Kubernetes-based production infrastructure
Improve the reliability, availability, and performance of Saturn Cloud services
Automate infrastructure provisioning, deployment, scaling, and operational workflows
Build and maintain infrastructure using tools such as Terraform and Helm
Develop tooling and services in Python and Go
Improve monitoring, logging, alerting, and observability across our systems
Investigate production issues, participate in incident response, and help prevent recurring failures
Identify operational bottlenecks and eliminate repetitive manual work through automation
Help evolve Saturn Cloud’s infrastructure across multiple cloud environments
Collaborate closely with engineering teams on architecture, reliability, scalability, and production readiness
What We’re Looking For
We’re looking for an engineer who is comfortable working deep in cloud infrastructure and production systems and who approaches reliability as an engineering problem rather than simply an operational responsibility.
Required
Strong production experience with Kubernetes
Experience with cloud infrastructure on AWS, GCP, Azure, or similar platforms
Experience with infrastructure-as-code, particularly Terraform
Experience deploying and managing Kubernetes applications with Helm
Experience programming in Python, Go, or both
Strong understanding of Linux, networking, containers, and distributed systems
Experience troubleshooting production systems and identifying root causes
Strong engineering practices around code quality, testing, maintainability, and automation
Ability to work effectively on a distributed, remote engineering team
Based in the APAC region
Nice to Have
Experience operating large or complex Kubernetes environments
Experience with observability tools such as Prometheus, Grafana, OpenTelemetry, or similar systems
Experience defining or working with SLIs, SLOs, and reliability metrics
Experience with multi-cloud infrastructure
Experience operating GPU or other compute-intensive workloads
Experience with AI/ML, data science, or distributed computing infrastructure
Experience working at an early-stage or high-growth technology company
Why Saturn Cloud
Remote-first culture with a high-trust, high-ownership environment
Work on challenging infrastructure problems at the center of modern AI and ML
Help shape the reliability and architecture of a rapidly evolving cloud platform
Significant opportunity to influence engineering practices and technical direction
Compensation & Benefits
Competitive salary
Flexible PTO
Fully remote position
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.