Live opening · Posted 12 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Site Reliability Engineer (SRE)
We are seeking Senior level Site Reliability Engineer (SRE) with strong AWS and automation experience to support disaster recovery, resiliency, and critical systems testing. This role will work closely with SRE engineers, technical leads, and architects to define, automate, execute, validate, and report on resilience and recovery testing across AWS environments. The ideal candidate will have 10+ years of SRE experience, strong Python coding skills, and extensive hands-on AWS experience. This is a highly technical team member role focused on engineering and execution rather than full software development lifecycle ownership.
Key Responsibilities
Design, automate, and execute reliability, resiliency, and disaster recovery testing across AWS environments.
Develop automated workflows for test execution, validation, results analysis, and reporting.
Partner with SRE engineers, technical leads, and architects to define effective testing strategies, patterns, and reusable approaches.
Use testing outcomes and lessons learned to help establish a Center of Excellence for enterprise resilience.
Evolve the organization's test portfolio to improve the resiliency of critical systems and applications.
Support disaster recovery testing involving large volumes of information and complex AWS environments.
Perform AWS fault injection and chaos testing to validate system behavior under failure conditions.
Develop automation and tooling using Python and AWS services.
Work extensively with AWS CLI and Linux-based environments to configure, troubleshoot, and validate infrastructure and applications.
Build and maintain infrastructure and testing automation using Terraform and AWS services including EC2, Lambda, and Auto Scaling.
Analyze test results, identify reliability and recovery gaps, and provide actionable recommendations.
Document testing patterns, procedures, results, and lessons learned to support knowledge sharing across the organization.
Help transition resilience and recovery testing into standardized enterprise testing practices.
Required Qualifications
5–10+ years of hands-on SRE experience in production environments.
Strong expertise with Amazon Web Services (AWS).
Strong Python programming and automation experience.
Extensive experience with AWS CLI and Linux.
Hands-on experience with Terraform and infrastructure automation.
Experience with AWS services such as EC2, Lambda, and Auto Scaling.
Experience designing and executing disaster recovery, resiliency, or reliability testing.
Experience working with critical systems where availability, resiliency, and recovery are important.
Ability to work effectively with engineers, technical leads, and architects to define and execute technical testing strategies.
Strong troubleshooting, analytical, and problem-solving skills.
Preferred Qualifications
Experience with AWS Fault Injection Service (FIS) or other chaos engineering/fault-injection platforms.
Experience developing automated chaos or resilience testing workflows.
AWS certifications are encouraged.
Experience establishing standardized testing practices, reusable patterns, or Centers of Excellence.
Experience with enterprise-scale AWS environments.
Experience supporting business continuity and disaster recovery initiatives.
Work Environment
Remote
AWS-focused environment
Highly technical SRE/engineering team
Citizenship requirement applies
Mid-to-senior level position; this is not intended for an entry-level or junior SRE.
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.