Live opening · Posted 6 hours ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Site Reliability Engineer I enhances system resilience and performance, implements automation tools, and contributes to the architectural design and disaster recovery strategies, promoting best practices for continuous improvement and reliability.
Education Qualifications:
Bachelor’s degree in Computer Science, Information Technology, Engineering, and/or comparable experience; advance degree preferred
Knowledge of modern observability stack – Splunk, Elastic Search, Prometheus, Grafana
Knowledge of containerization technologies (e.g., Kubernetes, Docker) and microservices architecture
Knowledge of observability tools and methodologies, including experience with logging, monitoring, tracing, and performance analysis platforms
Knowledge of cloud-based Site Reliability Engineering (SRE) practices and experience with public cloud platforms such as AWS, Azure, or Google Cloud
Work Experience:
Experience in software development, or technology operations, with a focus on Site Reliability Engineering
Experience in Linux/Unix systems, object-oriented programming languages (e.g., Java), scripting languages (e.g., Python, Bash), and cloud platforms (e.g., AWS, Azure, GCP)
Licenses and Certifications:
Advanced certification in Site Reliability Engineering (SRE) or related is a plus
Collaborates with Software Engineering teams to support the development, and implementation of features that enhance system resilience, scalability, and performance, ensuring systems can handle varying loads and recover effectively from unexpected disruptions
Collaborates in the development and implementation of automation tools and frameworks, including infrastructure as code (IaC) practices, to reduce manual intervention and improve system efficiency, with guidance from peers and leaders
Collaborates with senior engineers to contribute to the architectural design of systems, ensuring that reliability, scalability, and performance considerations are integrated into design discussions with direct guidance from senior colleagues
Collaborates in the design and execution of chaos engineering experiments and other resiliency testing methods to help identify potential failure points and ensure systems can recover from disruptions, with guidance from peers and leaders
Supports the development and implementation of disaster recovery plans and business continuity strategies, ensuring systems can recover quickly and effectively from unexpected disruptions
Collaborates with seniors to promote and implement best practices such as error budgeting, service-level objectives (SLOs), and service-level indicators (SLIs), contributing to a culture of continuous improvement and reliability
Collaborates and co-creates effectively with teams in product and the business to align technology initiatives with business objectives
Work arrangement
Hybrid
More openings worth a look
Recently tracked roles with full details and direct application links.