Live opening · Posted 5 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
We are looking for an enthusiastic and proactive Site Reliability Engineer to join our SRE team and help us ensure we provide world-class resilience and performance across the platform. The remit and focus of the role is to advise on all aspects of site reliability including availability, scalability, observability and capacity planning. It's a broad and exciting role, so we're looking for someone up for a challenge - if you're an energetic and a collaborative Site Reliability Engineer, this is the role for you.
Key responsibilities
Proactively monitor and analyse platform performance
Collaborate with engineering teams to address performance bottlenecks and ensure scalability
Assist engineering teams with implementing and reviewing SLOs
Continually improve observability through monitoring and alerting, and dashboards, using tools such as DataDog or Prometheus for example
Work with other teams to ensure it is effective and provides full coverage
Ensure the service is highly available and resilient
Champion best practices in design for high availability
Devise runbooks and run game sessions to test our DR plan, H/A and backups
Conduct assessments of capacity and plan for scaling to meet current and future business needs
Work closely with the Head of Platform Engineering and Head of SRE to strategize and implement scalable solutions
Work closely with the Platform team, feature teams and, 2nd line support and other stakeholders to ensure a good level of service is provided for our customers and embed SRE practices
Key player in the response and troubleshooting of incidents, ensuring rapid resolution and minimising downtime
Participate in blameless postmortems to identify root cause and corrective actions
Develop and maintain playbooks and documentation
Requirements
7-12 years of experience
Experience in performance monitoring and analysis
Capacity planning experience
Scripting and automation skills, with experience in relevant technologies
Experience with Infrastructure as Code, in particular, Terraform
Understanding of relational database technologies and their cloud versions (e.g. AWS Aurora)
Experience with messaging and distributed asynchronous workloads
Experience with nginx or similar technologies
Familiarity with SRE processes
Aware of DevOps principles like the 3 ways and 5 ideals
Desired Skills
Experience with other database technologies and cloud platforms
Past experience with Enterprise solutions running at scale
Familiarity with Kanban and Agile development processes
Experience with containerisation, for example Docker
Familiarity with software best practices such as Refactoring, Clean Code, Domain-Driven Design and Test-Driven Development
Benefits
The chance to work alongside a team of hard-working, passionate people in a role where you'll see the impact of your work everyday. We also offer:
Hybrid work environment
Group Term Life Insurance paid out at 3x Annual CTC (Arbor India)
32 days holiday (plus Arbor Holidays). This is made up of 25 days annual leave plus 7 extra companywide days given over Easter, Summer & Christmas
Work time: 9.30 am to 6 pm (8.5 hours only)
Compensation - 100% fixed salary disbursement and no variable components
Work arrangement
Hybrid
More openings worth a look
Recently tracked roles with full details and direct application links.