Live opening · Posted 7 days ago

Senior Site Reliability Engineer

asobbi · United States (Remote)
Linkedin Yes
You are 7 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 7 days ago
Companyasobbi
LocationUnited States (Remote)
Work modeYes
SourceLinkedin
Listed7 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
21 min from Linkedin publishing this role to us finding it
18 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
70,605 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

Senior Site Reliability / DevOps Engineer | US Remote | $150k–$190k + Bonus + Equity
Overview
We’re supporting a rapidly scaling Neocloud building large-scale GPU infrastructure for demanding AI training and inference workloads.
The platform combines GPU compute, high-performance networking, bare metal, Kubernetes and cloud infrastructure. The engineering team is deliberately small and senior, giving individuals genuine ownership rather than narrow responsibility.
They’re now hiring a Senior SRE / DevOps Engineer to take ownership of major parts of the production platform, improving reliability, automation and operational maturity as the infrastructure continues to scale.
The Opportunity
Own production infrastructure rather than simply supporting someone else’s platform.
Work directly with large-scale GPU and AI infrastructure.
Solve problems across Kubernetes, Linux, networking, storage and bare metal.
Join a senior engineering environment with significant technical autonomy.
Help bring new GPU clusters and infrastructure online.
Use modern automation and AI-assisted operational tooling to reduce manual work.
The Role
You’ll own the reliability and operability of key parts of the GPU cloud platform.
This is a hands-on senior IC role covering production engineering, incident response, automation, observability and infrastructure operations. You’ll also mentor less experienced engineers and participate in a shared on-call rotation.
Responsibilities
Own reliability and performance across major production platform components.
Lead technical response during production incidents and drive problems through to resolution.
Build automation and Infrastructure as Code to remove repetitive operational work.
Improve monitoring, alerting and observability across production infrastructure.
Debug complex Linux, networking, storage and performance issues.
Support the deployment and operational readiness of new infrastructure and GPU clusters.
Create practical runbooks and post-mortems that improve future operations.
Mentor engineers through reviews, pairing and incident response.
Skills & Experience
Essential
5+ years in SRE, DevOps, production engineering or infrastructure operations.
Strong hands-on Linux experience in production environments.
Deep operational Kubernetes experience, including troubleshooting at scale.
GPU or HPC infrastructure experience.
Good networking fundamentals across L2/L3 and BGP.
Strong Infrastructure as Code / automation experience using Terraform, Ansible or similar.
Experience owning on-call, incident response and post-mortems.
Comfortable taking end-to-end ownership of production systems.
Nice to Have
InfiniBand, NCCL or GPU monitoring experience.
OpenStack or bare-metal provisioning.
Prometheus, Grafana, Checkmk or similar observability tooling.
Previous experience within a Neocloud, cloud provider or infrastructure business.
Compensation
$150,000–$190,000 base salary, plus bonus, equity and benefits.
Interested?
Apply directly or message me for a confidential discussion to learn more.

Work arrangement
Yes

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App