Live opening · Posted 5 days ago

Senior Staff Software Engineer – SRE & AIOps

ServiceNow · Santa Clara, CALIFORNIA, United States
Smartrecruiters Hybrid Full-time
You are 5 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 5 days ago
CompanyServiceNow
LocationSanta Clara, CALIFORNIA, United States
Job typeFull-time
Work modeHybrid
SourceSmartrecruiters
Listed5 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
0 min from Smartrecruiters publishing this role to us finding it
1 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
16,433 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

About the Role
ServiceNow is seeking a Senior Staff Reliability Engineer – SRE & AIOps to drive infrastructure automation, operational resilience, and toil elimination across our hybrid cloud and data center operations. Embedded within the Site Reliability & Database Engineering organization, this technical leader will design and implement automation-first systems that reduce manual intervention, accelerate incident remediation, and enable our global engineering teams to operate reliably at scale.
This role combines deep hands-on technical expertise in Kubernetes, cloud platforms, and DevOps practices with strategic influence across infrastructure teams. You will architect SRE tooling, develop auto-remediation capabilities, and establish patterns that allow ServiceNow's cloud platform to maintain industry-leading reliability while minimizing operational toil across follow-the-sun global teams.
What you get to do in this role:
Design, deploy, and operate enterprise-scale Kubernetes clusters across hybrid and multi-cloud environments, establishing governance, scaling policies, and operational practices that support high-velocity application deployments at 99.99%+ availability targets.
Architect and implement closed-loop auto-remediation systems that detect, classify, and resolve transient infrastructure failures without human intervention, leveraging agentic AI and machine learning frameworks to predict failures, trigger preventive actions, and continuously reduce MTTR and on-call burden.
Design and evolve the SRE tooling stack, including monitoring platforms, incident management systems, log aggregation, and observability integrations, that support global follow-the-sun on-call operations and enable data-driven incident response.
Establish SLO frameworks, error budgets, and alerting policies that balance rapid incident response with alert fatigue management, while developing automated runbooks and playbooks that empower on-call engineers to resolve issues autonomously.
Design and maintain Infrastructure-as-Code frameworks and GitOps pipelines that enable reproducible, auditable infrastructure deployments across hybrid and multi-cloud environments with consistent security and compliance guardrails.
Architect hybrid cloud and data center operations, spanning on-premises infrastructure, public cloud environments, and edge computing, including workload migration strategies, disaster recovery patterns, and cost optimization practices across multi-region deployments.
Drive adoption of containerization, microservices, and DevOps patterns across engineering teams, establishing CI/CD best practices, service mesh architectures, and network security controls that enable rapid, safe release cycles.
Design on-call rotation schedules, escalation policies, and incident command systems that span across different time zones, ensuring 24/7 incident response while driving post-incident review processes that capture learning and drive systemic improvements.
Mentor and guide junior SRE engineers, infrastructure teams, and DevOps practitioners on advanced reliability patterns, incident investigation techniques, automation best practices, and agentic AI applications for infrastructure operations.
Champion a culture of blameless incident analysis, data-driven decision-making, continuous improvement, and experimentation across engineering teams, establishing knowledge-sharing practices and technical documentation standards.
Reduce operational toil through systematic automation of repetitive tasks, from infrastructure provisioning to incident response to cost optimization, directly improving team capacity and job satisfaction across globally distributed operations.
To be successful in this role you have:
Kubernetes Mastery: Deep, hands-on expertise operating production Kubernetes clusters at scale, including cluster design, node management, pod orchestration, resource quotas, network policies, security controls, and troubleshooting complex runtime issues.
Incident Auto-Remediation Expertise: Proven experience designing and implementing closed-loop automated remediation systems, including anomaly detection, alert correlation, runbook automation, and self-healing mechanisms, that measurably reduce MTTR and on-call burden.
Cloud Platform Experience: Extensive hands-on experience with AWS (EKS, EC2, RDS, Lambda), Azure (AKS, VMs, CosmosDB), and GCP (GKE, Compute Engine, Cloud SQL), capable of architecting multi-region, multi-cloud solutions.
DevOps & IaC Proficiency: Expert-level experience with Infrastructure-as-Code tools and GitOps platforms to drive reproducible, auditable infrastructure deployments.
SRE Tooling Fluency: Strong working knowledge of observability platforms, incident management systems, and log aggregation.
Distributed Systems Thinking: Deep understanding of distributed system challenges, eventual consistency, cascading failures, network partitions, Byzantine fault tolerance, and proven ability to design systems resilient to these conditions.
On-Call Operations: Experience operating in follow-the-sun, 24/7 on-call models; ability to design escalation policies, runbooks, and communication patterns that balance responsiveness with operator well-being.
Data Center & Hybrid Cloud Operations: Hands-on experience managing both on-premises infrastructure and public cloud environments, including hybrid networking, disaster recovery, and workload migration strategies.
AI/ML Integration: Demonstrated ability to apply machine learning and AI-driven insights to infrastructure operations, including anomaly detection, predictive alerting, and intelligent remediation.
Influence Without Authority: Proven ability to drive technical decisions across teams, influence architecture choices, and mentor engineers outside direct reporting structure through credibility and technical depth.
Qualifications
Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving. This may include using AI-powered tools, automating workflows, analyzing AI-driven insights, or exploring AI's potential impact on the function or industry.
12+ years in software engineering or infrastructure operations, with 7+ years in senior SRE, DevOps, or cloud platform engineering roles managing large-scale distributed systems with a Bachelor's degree; or 8 years and a Master's degree; or a PhD with 5 years experience; or equivalent experience.
5+ years hands-on experience designing, deploying, and operating production Kubernetes clusters at enterprise scale.
Proficiency in Infrastructure-as-Code: Terraform, CloudFormation, or equivalent tools used to manage infrastructure at scale.
Public Cloud Expertise: Demonstrable experience across 2+ of the following: AWS, Azure, GCP, with deep knowledge of services relevant to SRE operations (compute, networking, storage, observability).
On-Call Leadership: Experience designing or operating 24/7 follow-the-sun on-call models for globally distributed teams, including escalation policies, runbook development, and incident response.
Incident Auto-Remediation: Proven ability to architect and implement automated remediation systems, from basic alert automation to sophisticated closed-loop AI-driven systems, that meaningfully reduce manual toil.
Linux & Systems Programming: Strong foundation in Linux system administration, performance troubleshooting, and scripting (Python, Go, Bash).
SRE Mindset: Demonstrated commitment to reliability through engineering, favoring durable automation over heroics, data-driven decision-making, and continuous learning.
Bachelor’s degree in computer science, Computer Engineering, or related field (or equivalent professional experience).
Preferred:
Kubernetes certification (CKA, CKAD, or equivalent).
AI/ML certification or demonstrated expertise in agentic AI frameworks and machine learning applications for infrastructure operations.
Experience with service mesh platforms or advanced networking in Kubernetes environments.
Background in migrating workloads from on-premises data centers to public cloud environments.
Experience with cost optimization practices in hybrid cloud environments (reserved instances, spot instances, resource right-sizing).
Track record of mentoring or leading infrastructure engineering teams.
Why This Role?
This role offers the opportunity to eliminate operational toil at scale and build the automation-first infrastructure practices that define modern cloud operations. You will architect systems that allow ServiceNow's global teams to sleep soundly on-call, confident that auto-remediation systems are actively preventing and resolving failures. Your work will directly shape how the organization scales reliability as it grows, establishing patterns that persist long after you've architected them. This is a role for an engineer who wants technical depth, strategic influence, and the satisfaction of watching a system recover from failure without waking anyone up.
For positions in

Employment type
Full-time

Work arrangement
Hybrid

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App