Live opening · Posted 7 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Senior Site Reliability / DevOps Engineer | US Remote | $150k–$190k + Bonus + Equity
Overview
We’re supporting a rapidly scaling Neocloud building large-scale GPU infrastructure for demanding AI training and inference workloads.
The platform combines GPU compute, high-performance networking, bare metal, Kubernetes and cloud infrastructure. The engineering team is deliberately small and senior, giving individuals genuine ownership rather than narrow responsibility.
They’re now hiring a Senior SRE / DevOps Engineer to take ownership of major parts of the production platform, improving reliability, automation and operational maturity as the infrastructure continues to scale.
The Opportunity
Own production infrastructure rather than simply supporting someone else’s platform.
Work directly with large-scale GPU and AI infrastructure.
Solve problems across Kubernetes, Linux, networking, storage and bare metal.
Join a senior engineering environment with significant technical autonomy.
Help bring new GPU clusters and infrastructure online.
Use modern automation and AI-assisted operational tooling to reduce manual work.
The Role
You’ll own the reliability and operability of key parts of the GPU cloud platform.
This is a hands-on senior IC role covering production engineering, incident response, automation, observability and infrastructure operations. You’ll also mentor less experienced engineers and participate in a shared on-call rotation.
Responsibilities
Own reliability and performance across major production platform components.
Lead technical response during production incidents and drive problems through to resolution.
Build automation and Infrastructure as Code to remove repetitive operational work.
Improve monitoring, alerting and observability across production infrastructure.
Debug complex Linux, networking, storage and performance issues.
Support the deployment and operational readiness of new infrastructure and GPU clusters.
Create practical runbooks and post-mortems that improve future operations.
Mentor engineers through reviews, pairing and incident response.
Skills & Experience
Essential
5+ years in SRE, DevOps, production engineering or infrastructure operations.
Strong hands-on Linux experience in production environments.
Deep operational Kubernetes experience, including troubleshooting at scale.
GPU or HPC infrastructure experience.
Good networking fundamentals across L2/L3 and BGP.
Strong Infrastructure as Code / automation experience using Terraform, Ansible or similar.
Experience owning on-call, incident response and post-mortems.
Comfortable taking end-to-end ownership of production systems.
Nice to Have
InfiniBand, NCCL or GPU monitoring experience.
OpenStack or bare-metal provisioning.
Prometheus, Grafana, Checkmk or similar observability tooling.
Previous experience within a Neocloud, cloud provider or infrastructure business.
Compensation
$150,000–$190,000 base salary, plus bonus, equity and benefits.
Interested?
Apply directly or message me for a confidential discussion to learn more.
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.