Live opening · Posted 27 days ago

Lead Site Reliability Engineer

Zeta · Hyderabad
Instahyre 12-14 yrs
You are 27 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 27 days ago
CompanyZeta
LocationHyderabad
Experience12-14 yrs
SourceInstahyre
Listed27 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
0 min from Instahyre publishing this role to us finding it
8 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
15,816 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

Responsibilities:
System Reliability: Ensuring the reliability of software systems by designing, implementing, and maintaining scalable and reliable infrastructure.
Automation: Developing automation tools and scripts to streamline operational tasks, reduce manual intervention, and improve overall system efficiency.
Incident Response and Resolution: Monitoring system performance and responding to incidents promptly to minimise downtime and ensure high availability.
Capacity Planning: Analysing system usage patterns and forecasting future capacity needs to ensure that the infrastructure can handle current and future demands.
Performance Optimisation: Identifying and addressing performance bottlenecks in software systems through optimisation and tuning.
Infrastructure as Code (IaC): Implementing infrastructure as code practices, using tools like Terraform or Ansible, to define and manage infrastructure in a version-controlled and automated manner.
Monitoring and Logging: Implementing and maintaining monitoring and logging solutions to gain insights into system behaviour, troubleshoot issues, and proactively address potential problems.
Security: Collaborating with security teams to implement and maintain security best practices in infrastructure and applications.
Disaster Recovery Planning: Developing and maintaining disaster recovery plans to ensure that systems can quickly recover from major outages or failures.
Continuous Improvement: Continuously analysing system performance, reliability, and incidents to identify areas for improvement and implementing changes to enhance overall system resilience.
Team Leadership: Ability to lead and motivate a team of SREs, providing guidance and support.
Mentorship and Coaching: Providing mentorship and coaching to team members to foster their professional development.
Conflict Resolution: Skill in resolving conflicts and addressing challenges within the team.
Requirements:
Programming Languages: Proficiency in one or more programming languages, commonly Python, Go, Shell, and Bash.
Automation and Scripting: Strong automation skills using tools like Ansible, Puppet, Chef, or custom scripts. Knowledge of Infrastructure as Code (IaC) tools like Terraform.
Containerisation and Orchestration: Experience with containerisation technologies like Docker and container orchestration platforms like Kubernetes.
Cloud Computing: Proficiency in any of the cloud platforms, such as AWS, Azure, or Google Cloud Platform, and knowledge of managing infrastructure in the cloud.
Monitoring and Logging: Familiarity with monitoring tools (e. g., Prometheus, Grafana, and ELK Stack) and logging frameworks to track system performance and troubleshoot issues.
Networking: Understanding of networking concepts, protocols, and troubleshooting skills.
Security: Knowledge of security best practices, including encryption, access controls, and vulnerability management.
Continuous Integration/Continuous Deployment (CI/CD): Understanding and implementation of CI/CD pipelines for automated testing and deployment.
Load Balancing: Experience in incident response, troubleshooting, and resolution.
Version Control: Proficient use of version control systems like Git.
12 - 14 years of experience in site reliability engineering.
B. Tech/M. Tech in computer science, information technology, or a related field.
Having experience working for a product organisation is a plus.
Certifications from cloud service providers like AWS Certified DevOps Engineer, Google Cloud Professional DevOps Engineer, or Microsoft Certified are a plus.

Experience
12-14 yrs

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App