Live opening · Posted 8 days ago

Staff Reliability Engineer

ServiceNow · Santa Clara, CALIFORNIA, United States
Smartrecruiters Yes Full-time
You are 8 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 8 days ago
CompanyServiceNow
LocationSanta Clara, CALIFORNIA, United States
Job typeFull-time
Work modeYes
SourceSmartrecruiters
Listed8 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
8 min from Smartrecruiters publishing this role to us finding it
11 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
15,760 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

Join us to build the next generation of cloud-native reliability, release, and test platforms that enable engineering excellence, developer productivity, and high-confidence ServiceNow releases through automation, observability, and AI-driven operations.
What you get to do in this role:
Design, build, and operate cloud-native engineering platforms for software validation, release validation, and production readiness
Design and maintain production-like release and test ServiceNow environments that improve release confidence and deployment readiness.
Build and integrate automated test pipelines, observability, reliability signals, deployment intelligence, and quality gates into CI/CD workflows.
Develop automation solutions that improve engineering productivity, streamline operations, and reduce manual toil through shift-left engineering practices.
Build reusable frameworks, self-service engineering environments, test data management, mock services, and developer productivity tooling.
Design and enhance Kubernetes-based platforms supporting scalable test infrastructure, release automation, cloud-native workloads, and developer self-service.
Implement automated validation for failure detection, deployment verification, policy enforcement, security checks, resilience testing, and operational health assessments.
Resolve complex platforms, infrastructure, and networking challenges through software engineering, systems design, and automation.
Partner closely with engineering teams to improve platform reliability, release quality, cloud-native adoption, and engineering best practices.
Participate in architecture reviews, technical design discussions, and implementation of scalable, automation-first engineering solutions.
Influence technical decisions through strong engineering execution, collaboration, and delivery of high-quality platform capabilities.
Mentor engineers through technical guidance, code reviews, knowledge sharing, and engineering best practices.
Foster a culture of reliability, automation, operational excellence, continuous improvement, and customer-focused engineering.
To be successful in this role you have:
Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving. This may include using AI-powered tools, automating workflows, analyzing AI-driven insights, or exploring AI's potential impact on the function or industry.
8+ years of experience in Site Reliability Engineering (SRE), DevOps, Platform Engineering, Software Engineering, or Infrastructure Engineering with a Bachelor's degree; or 6 years and a Master's degree; or a PhD with 3 years experience; or equivalent experience.
Hands-on experience with Kubernetes across cluster operations, networking, storage, security, autoscaling, and multi-cluster environments.
Experience building and operating cloud-native platforms supporting scalable, highly available services.
Experience integrating Kubernetes with CI/CD, GitOps, automated test pipelines, deployment validation, and cloud-native deployment workflows.
Experience designing and implementing automation to improve developer productivity, release quality, and operational efficiency.
Experience with progressive delivery practices, including canary deployments, feature flags, automated rollback, and deployment verification.
Experience with chaos engineering, resilience testing, disaster recovery, and reliability validation.
Strong software engineering skills with hands-on experience designing, developing, testing, and debugging applications using Python, Go, Java, or Ruby.
Experience leveraging AI-assisted engineering for intelligent testing, release risk analysis, incident diagnostics, or operational automation is a plus.
Strong understanding of observability, monitoring, SLI/SLOs, incident management, and production operations for distributed systems.
Demonstrated ability to solve complex technical problems, drive projects independently, and collaborate effectively across engineering teams.
Thrives in fast-paced, ambiguous environments with a strong ownership mindset, bias for action, and a passion for continuous learning and automation.
Low ego, intellectually curious, and an effective collaborator who enjoys partnering with globally distributed teams to deliver reliable engineering solutions.
Good to have:
Experience with observability and monitoring platforms for applications, services, and distributed systems at scale.
Experience with DevOps automation, CI/CD pipelines, GitOps, and Agile development practices using tools such as GitLab CI/CD, Argo CD, or Flux.
Experience building and maintaining enterprise-scale test automation frameworks using technologies such as Playwright, Selenium, Cypress, REST Assured, PyTest, JUnit/TestNG, or equivalent.
Experience with test orchestration, intelligent regression testing, test impact analysis, flaky test detection, parallel execution, and test data management.
Experience with service virtualization, contract testing, synthetic testing, and building developer self-service engineering platforms.
Experience with Infrastructure as Code and configuration management tools such as Ansible, Terraform, or equivalent.
Experience with the Kubernetes ecosystem, including Helm, Argo Workflows, Kustomize, Istio/Linkerd, Gateway API/Ingress, Prometheus, OpenTelemetry, and container runtime technologies.
Experience operating Kubernetes platforms across public cloud providers, including AWS (EKS), Azure (AKS), and Google Cloud (GKE).
Experience implementing progressive delivery practices, including canary deployments, feature flags, deployment verification, and automated rollback.
Familiarity with AI-assisted engineering, intelligent testing, operational automation, or cloud-native engineering platforms.
For positions in this location, we offer a base pay of $166,500 - $291,400, plus equity (when applicable), variable/incentive compensation and benefits. Sales positions generally offer a competitive On Target Earnings (OTE) incentive compensation structure. Please note that the base pay shown is a guideline, and individual total compensation will vary based on factors such as qualifications, skill level, competencies, and work location. We also offer health plans, including flexible spending accounts, a 401(k) Plan with company match, ESPP, matching donations, a flexible time away plan and family leave programs. Compensation is based on the geographic location in which the role is located and is subject to change based on work location.

Employment type
Full-time

Work arrangement
Yes

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App