Live opening · Posted 5 days ago

Staff Site Reliability Engineer

Doghouse Recruitment · United States (Remote)
Linkedin Yes
You are 5 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 5 days ago
CompanyDoghouse Recruitment
LocationUnited States (Remote)
Salary5 benefits
Work modeYes
SourceLinkedin
Listed5 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
21 min from Linkedin publishing this role to us finding it
4 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
72,525 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

US - 100% Remote.
Staff Site Reliability Engineer – Bare Metal Linux – Data Center – Networking
OTE up to $350k (base + variable), depending on experience
Our client is building a cloud platform for high-throughput, compute-heavy workloads. They operate large-scale infrastructure where failure modes are real, capacity is finite, and reliability needs to be engineered, not "handled".
We're seeking a Senior/Staff SRE who will own production reliability end-to-end for our client: define SLIs/SLOs, run error budget conversations, and ship changes that reduce incidents and improve latency (p95/p99). You'll build automation to kill toil, improve deployment safety (canary/rollback), and turn observability into signal rather than noise.
This is a bare-metal environment: think Linux, datacenters, physical fleets, and real hardware constraints, not managed services. You'll work close to the metal across Kubernetes internals (scheduling, autoscaling behavior, kubelet pressure/evictions, etcd/control plane), Linux performance (CPU/memory/IO contention), and network debugging (DNS/TCP/TLS, packet loss, congestion). On-call is part of the job, but success is measured by how much you reduce it.
Must requirements:
Extensive Production Engineering experience running bare metal / on-prem / data center infrastructure (not public cloud only)
Deep hands-on expertise in Linux systems debugging and performance (CPU, memory, IO, -level behaviors)
Strong understanding of networking (DNS/TCP/TLS, latency, packet loss, congestion, troubleshooting underload)
Strong Kubernetes experience beyond manifests: scheduler behavior, autoscaling edge cases, kubelet pressure/evictions, etcd/control plane
Experience with Terraform, Docker, Helm, and modern CI/CD practices
Strong coding skills are required for this role either in Go, and/or Python, beyond automation scripting - Real engineering capability is a must
Experience in Low Latency environments.
If you're looking for complexity and a new place to nerd out on infrastructure optimization, we'd love to hear fro
Location: United States - 100% Remote.
Total compensation: OTE up to $350k (base + variable), depending on experience

Work arrangement
Yes

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App