Live opening · Posted 3 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Job description:
Job Description
Job Title Cloud Site Reliability Engineer SRE
Exp : 5 to 12 years
Position Overview
We are seeking a Cloud Site Reliability Engineer SRE to drive the reliability scalability and performance of our cloudbased infrastructure The ideal candidate combines software engineering expertise with advanced systems operations skills to maintain highly available systems while reducing operational toil This role involves automation monitoring capacity planning incident response and cloud platform management across a dynamic distributed environment
As a Cloud SRE you will work closely with Engineering Architecture DevOps and security teams to ensure seamless service experiences for our customers while contributing to platform design and operational efficiency
Position Requirements
Our Engineers play a critical role in the success of our clients and are expected to effectively communicate our recommended solutions in a consultative role for each client Therefore a successful candidate will possess a high degree of selfmanagement personal accountability strong communication skills and teamwork The ability to interact engineer and communicate collaboratively at the highest technical levels with customers vendors partners and all members of staff is required
Monitoring Observability Implement monitoring ing and logging frameworks using Splunk Azure monitor Dynatrace AWS cloud watch or similar to detect and resolve issues proactively
Required Skills Qualifications
Programming Scripting Proficiency in Python Power Shell Bash or equivalent for automation and system management
Cloud Platforms Handson experience with AWS Azure or GCP strong understanding of VPCs IAM serverless architectures and managed Kubernetes services
Containers Orchestration Experience with Docker and Kubernetes
Infrastructure as Code IaC Proficient in Terraform Ansible
Monitoring Observability Expertise with Splunk Azure Monitor Dynatrace AWS Cloud Watch or similar tools
Expert Knowledge and practical experience using Cloud data migration tools
Operating Systems Advanced knowledge of Windows LinuxUnix environments with experience in system administration and networking fundamentals
Incident Response Strong problemsolving skills under pressure with experience managing outages and mitigating risk
Collaboration Communication Ability to articulate technical insights coordinate across teams and contribute to a blameless culture to resolve issues and drive consistent results
Preferred Qualifications
Industry certifications such as AWS Certified Solutions Architect Google Cloud Professional DevOps Engineer Azure Dev Ops Engineer
Exposure to chaos engineering or resilience testing frameworks
Prior experience in multicloud deployments or hybrid cloud environments
Familiarity with servicelevel objectives SLOs indicators SLIs and error budgets for service reliability
Gather feedback from the department on areas of improvement and provide solutions utilizing Azure
Skills:
Mandatory Skills : Automation & Scripting
Good to Have Skills : Critical Incident Response, Monitoring & Observability, Service Level & Error Budget Management
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.