Live opening · Posted 8 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
The core responsibilities for the job include the following:
Reliability and Platform Engineering:
Design and implement highly available, scalable, and fault-tolerant cloud infrastructure.
Establish and drive SRE best practices, including SLI, SLO, SLA, error budgets, and reliability reviews.
Lead architecture decisions for platform scalability, resilience, disaster recovery, and business continuity.
Improve system performance, availability, and operational efficiency across production environments.
Infrastructure as Code and Automation:
Build and maintain infrastructure using Terraform and cloud-native automation frameworks.
Develop automation solutions using Python, Bash, and PowerShell.
Drive infrastructure standardization and self-service platform capabilities.
Automate provisioning, deployments, security controls, and operational workflows.
Kubernetes and Cloud Operations:
Design, deploy, and manage large-scale Kubernetes environments.
Implement Helm-based deployment strategies and Kubernetes operational best practices.
Optimize cluster performance, security, capacity planning, and resource utilization.
Lead container platform modernization initiatives.
CI/CD and Developer Productivity:
Build and enhance CI/CD pipelines using Jenkins, GitHub Actions, and GitLab CI/CD.
Improve deployment reliability, release automation, and engineering productivity.
Enable DevSecOps practices across the software delivery lifecycle.
Observability and Incident Management:
Design and maintain observability platforms using Prometheus and Grafana.
Establish proactive monitoring, alerting, logging, and incident response frameworks.
Lead root cause analysis (RCA), postmortems, and continuous improvement initiatives.
Reduce operational toil through automation and intelligent alerting.
Security and Compliance:
Implement secure infrastructure practices across cloud and Kubernetes environments.
Manage IAM frameworks, secrets management, and privileged access controls.
Drive adoption of HashiCorp Vault for enterprise-grade secrets management.
Partner with Security teams to ensure compliance and governance requirements are met.
Leadership and Mentorship:
Act as a technical leader and trusted advisor across engineering teams.
Mentor SREs, DevOps Engineers, and Platform Engineers.
Drive engineering excellence through architecture reviews, technical guidance, and operational best practices.
Influence long-term platform strategy and infrastructure roadmap.
Requirements:
10+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Infrastructure Engineering.
Proven experience supporting enterprise-scale SaaS platforms.
Strong background working within product-based organizations.
Experience managing production environments with high availability and uptime requirements.
Demonstrated success in driving reliability initiatives across distributed systems.
Experience handling large-scale incidents, capacity planning, and production operations.
CI/CD and DevOps: Jenkins, GitHub Actions, GitLab CI/CD, Release automation and deployment orchestration.
Observability: Prometheus, Grafana, Monitoring, alerting, and telemetry design.
Configuration Management: Ansible, Chef, Puppet.
Security: IAM, HashiCorp Vault, Security automation and secrets management.
Databases: PostgreSQL, MySQL. Performance tuning, backup/recovery, and operational management experience.
Cloud and Infrastructure:
Strong experience with AWS cloud services and cloud-native architectures.
Deep expertise in Infrastructure as Code (Terraform).
Experience designing and operating large-scale distributed systems.
Containers and Orchestration:
Extensive hands-on experience with Kubernetes.
Strong expertise with Helm and containerized application deployment strategies.
Programming and Automation:
Strong coding skills in Python.
Proficiency in Bash and/or PowerShell scripting.
Experience building operational tooling and automation frameworks.
Experience
9-13 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.