Live opening · Posted 3 days ago

Site Reliability Engineer

Evlo AI · San Diego, CA (Remote)
Linkedin Yes
You are 3 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 3 days ago
CompanyEvlo AI
LocationSan Diego, CA (Remote)
Work modeYes
SourceLinkedin
Listed3 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
26 min from Linkedin publishing this role to us finding it
11 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
61,009 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

About The Role
The Site Reliability Engineer owns the systems and practices that keep production services available, performant, and recoverable at scale. The role focuses on Kubernetes-based workloads, cloud infrastructure, observability, incident response, and automation across distributed systems.
This engineer will work with platform, software, and security teams to reduce operational risk and improve the developer experience. The work includes defining service-level objectives, eliminating recurring failure modes, and building reliable delivery and recovery processes for customer-facing systems.
Key Responsibilities
Design and operate highly available infrastructure on AWS or GCP using Kubernetes, Terraform, and infrastructure-as-code standards
Build and maintain observability across services using Prometheus, Grafana, OpenTelemetry, and centralized logging platforms such as ELK or Datadog
Define and track SLIs, SLOs, error budgets, and capacity plans for critical production services
Automate deployment, scaling, backup, and recovery workflows through CI/CD pipelines using tools such as GitHub Actions, GitLab CI, or Argo CD
Lead incident response, including triage, communications, root-cause analysis, and durable remediation of recurring production issues
Harden systems through access controls, secrets management, patching, disaster recovery testing, and infrastructure reliability reviews
Partner with application teams to improve service design, operational readiness, performance, and on-call effectiveness
What We Are Looking For
3–8 years of experience in site reliability engineering, DevOps, platform engineering, or production infrastructure roles
Hands-on experience operating Kubernetes and containerized services in AWS, GCP, or Azure
Strong Linux administration and networking fundamentals, including DNS, HTTP, TLS, TCP/IP, load balancing, and troubleshooting distributed systems
Proficiency with Terraform or an equivalent infrastructure-as-code tool, plus practical experience designing CI/CD pipelines
Experience implementing observability with metrics, logs, traces, dashboards, and actionable alerting using tools such as Prometheus, Grafana, OpenTelemetry, or Datadog
Strong scripting or programming skills in Python, Go, or Bash, with a focus on automation, testing, and maintainable operational tooling
Bonus: Experience with service meshes, Kafka, PostgreSQL or Redis operations, compliance-focused infrastructure, chaos engineering, or a bachelor's degree in computer science, engineering, or a related field

Work arrangement
Yes

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App