Live opening · Posted 2 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Join a team focused on keeping cloud-based digital commerce services reliable, secure, and high-performing. This Sr. Site Reliability Engineer role offers the opportunity to strengthen monitoring, automation, and incident response across Kubernetes-based applications running in Azure.
Responsibilities
Support the reliability, availability, performance, and efficiency of cloud services and supporting infrastructure
Design, implement, and improve Site Reliability Engineering practices for cloud products and services
Build and enhance monitoring that detects symptoms early and helps prevent service disruptions
Monitor and troubleshoot Kubernetes-based applications and services running in Azure Kubernetes Service (AKS)
Collaborate with DevOps, security, architecture, infrastructure, network, and development teams to resolve cross-functional issues
Analyze key performance indicators and telemetry to identify trends, bottlenecks, and areas for improvement
Automate and strengthen operational processes to improve resilience, scalability, and security
Document processes, findings, and supporting materials to improve clarity and team knowledge sharing
Participate in compliance and regulatory activities as needed
Skills
5+ years of experience in software engineering, operations engineering, or a related technical role
2+ years of experience in DevOps, Site Reliability Engineering, or a similar cloud-native environment
Hands-on experience supporting cloud-based web applications in Microsoft Azure
Strong knowledge of monitoring infrastructure, application uptime, latency, and performance in distributed systems
Experience building or improving CI/CD pipelines
Solid troubleshooting skills across cloud and systems environments
Working knowledge of systems, storage, networking, security, and databases
Experience with version control tools such as Git, SVN, or CVS
Excellent written and verbal communication skills
Ability to collaborate effectively across technical and non-technical teams
Preferred Skills
Experience with observability and monitoring tools in Kubernetes environments, especially AKS
Familiarity with monitoring solutions such as Dynatrace, Azure Monitor, and Application Insights
Experience creating alerting strategies based on proactive, symptom-driven thresholds
Background in performance analysis using telemetry and monitoring data
Bachelor’s degree in Computer Science, Management Information Systems, or a related field, or equivalent experience
A proactive mindset focused on continuous improvement and service reliability
Horizontal is committed to building an inclusive workplace where different backgrounds, perspectives, and experiences are valued. We encourage qualified candidates to apply and join a team that believes equity, respect, and collaboration lead to better outcomes for everyone.
By applying for this position, you acknowledge and agree that Horizontal Talent may contact you regarding your application using automated technology, including phone calls, SMS/text messages, or email, which may be delivered by our virtual AI recruiter, Alex.
Please apply through this online posting or by visiting our Job Board at www.horizontaltalent.com/job-board. Applications will be accepted for 4 weeks. For those that join the team, we offer competitive compensation and benefits including medical, dental, vision, and retirement. Check out all we have to offer and how you can become part of the Horizontal Talent Team. The pay range for this role is $67 - $103 per hour based on qualifications and experience.
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.