Live opening · Posted 8 days ago

Principal Site Reliability Engineer

Swimlane · Work From Home
Instahyre 9-13 yrs
You are 8 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 8 days ago
CompanySwimlane
LocationWork From Home
Experience9-13 yrs
SourceInstahyre
Listed8 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
0 min from Instahyre publishing this role to us finding it
11 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
71,130 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

The core responsibilities for the job include the following:
Reliability and Platform Engineering:
Design and implement highly available, scalable, and fault-tolerant cloud infrastructure.
Establish and drive SRE best practices, including SLI, SLO, SLA, error budgets, and reliability reviews.
Lead architecture decisions for platform scalability, resilience, disaster recovery, and business continuity.
Improve system performance, availability, and operational efficiency across production environments.
Infrastructure as Code and Automation:
Build and maintain infrastructure using Terraform and cloud-native automation frameworks.
Develop automation solutions using Python, Bash, and PowerShell.
Drive infrastructure standardization and self-service platform capabilities.
Automate provisioning, deployments, security controls, and operational workflows.
Kubernetes and Cloud Operations:
Design, deploy, and manage large-scale Kubernetes environments.
Implement Helm-based deployment strategies and Kubernetes operational best practices.
Optimize cluster performance, security, capacity planning, and resource utilization.
Lead container platform modernization initiatives.
CI/CD and Developer Productivity:
Build and enhance CI/CD pipelines using Jenkins, GitHub Actions, and GitLab CI/CD.
Improve deployment reliability, release automation, and engineering productivity.
Enable DevSecOps practices across the software delivery lifecycle.
Observability and Incident Management:
Design and maintain observability platforms using Prometheus and Grafana.
Establish proactive monitoring, alerting, logging, and incident response frameworks.
Lead root cause analysis (RCA), postmortems, and continuous improvement initiatives.
Reduce operational toil through automation and intelligent alerting.
Security and Compliance:
Implement secure infrastructure practices across cloud and Kubernetes environments.
Manage IAM frameworks, secrets management, and privileged access controls.
Drive adoption of HashiCorp Vault for enterprise-grade secrets management.
Partner with Security teams to ensure compliance and governance requirements are met.
Leadership and Mentorship:
Act as a technical leader and trusted advisor across engineering teams.
Mentor SREs, DevOps Engineers, and Platform Engineers.
Drive engineering excellence through architecture reviews, technical guidance, and operational best practices.
Influence long-term platform strategy and infrastructure roadmap.
Requirements:
10+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Infrastructure Engineering.
Proven experience supporting enterprise-scale SaaS platforms.
Strong background working within product-based organizations.
Experience managing production environments with high availability and uptime requirements.
Demonstrated success in driving reliability initiatives across distributed systems.
Experience handling large-scale incidents, capacity planning, and production operations.
CI/CD and DevOps: Jenkins, GitHub Actions, GitLab CI/CD, Release automation and deployment orchestration.
Observability: Prometheus, Grafana, Monitoring, alerting, and telemetry design.
Configuration Management: Ansible, Chef, Puppet.
Security: IAM, HashiCorp Vault, Security automation and secrets management.
Databases: PostgreSQL, MySQL. Performance tuning, backup/recovery, and operational management experience.
Cloud and Infrastructure:
Strong experience with AWS cloud services and cloud-native architectures.
Deep expertise in Infrastructure as Code (Terraform).
Experience designing and operating large-scale distributed systems.
Containers and Orchestration:
Extensive hands-on experience with Kubernetes.
Strong expertise with Helm and containerized application deployment strategies.
Programming and Automation:
Strong coding skills in Python.
Proficiency in Bash and/or PowerShell scripting.
Experience building operational tooling and automation frameworks.

Experience
9-13 yrs

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App