Live opening · Posted 27 days ago

Site Reliability Engineer 3

PhonePe · Bangalore
Instahyre 7-11 yrs
You are 27 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 27 days ago
CompanyPhonePe
LocationBangalore
Experience7-11 yrs
SourceInstahyre
Listed27 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
0 min from Instahyre publishing this role to us finding it
11 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
15,760 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

We are seeking a highly motivated and experienced Site Reliability Engineer (SRE) with 7 to 12 years of experience to manage, scale, and ensure the high availability of our core infrastructure. This role involves deep expertise in cloud services, automation, monitoring, and complex networking to support a high-volume, mission-critical environment.
Responsibilities:
Cloud and Infrastructure: Configure, maintain, and manage services and packages on Ubuntu Virtual Machines in Azure. Design and manage Azure components for log storage, management, alerting, and monitoring.
Networking and Connectivity: Configure and maintain complex network components, including Azure Firewall, Route Tables, Virtual Network Gateways, and Express Route. Establish and manage IPsec and Express Route connectivity with external environments. Manage routing, troubleshooting connectivity issues, and support network component migrations with minimal downtime.
Automation and IaC: Drive automation for all BAU tasks using Terraform, SaltStack, Ansible, and scripting languages. Write new Terraform code for infrastructure components.
Database and Data Management: Set up and manage high-availability services like Mysql and Aerospike. Implement database replication across regions, manage migrations, and ensure data sync. Handle backups of databases, logs, and configurations.
Monitoring and Observability: Implement and manage monitoring (e. g., Prometheus, Victoria Metrics, Riemann) and centralised logging (Loki) solutions, with visualisation on Grafana. Troubleshoot performance and system issues at the OS, platform, or application level.
Security and Compliance: Manage firewalls and integrate platform and VM-level services with the SOC. Collaborate with Infosec teams to evaluate and fix security vulnerabilities.
Capacity and Performance: Conduct proactive capacity planning. Manage critical infrastructure components like Nginx, HA Proxy, Docker, and RMQ.
Incident Management and DR: Participate in an on-call rotation. Structure and lead incident response, Root Cause Analysis (RCA), and post-mortem creation. Set up and support the planning and execution of DR sites and failovers.
Requirements:
Core Services: Deep, hands-on experience with Microsoft Azure components, including Virtual Machines (Ubuntu/Linux), Azure Storage Accounts, CosmosDB, and Azure Data Explorer (ADX).
Networking: Expert-level knowledge in configuring and managing complex Azure networking components: Azure Firewall, Azure Route Tables, Virtual Network Gateways, Azure Express Route, and Azure Private DNS. Must be proficient in setting up and troubleshooting routing using protocols like BGP with on-prem DCs and managing network component migrations with minimal downtime.
Security/Compliance: Experience integrating platform and VM-level services with the Security Operations Centre (SOC) and collaborating with Infosec teams on vulnerability evaluation and remediation.
OS: Expert proficiency in Linux environments, specifically Ubuntu/Linux, for system administration, service configuration, and performance troubleshooting at the OS level.
High-Level Language: Deep expertise in at least one high-level language (Python, Go, or Java) for writing automation, services, and tooling.
Shell Scripting: Shell scripting (Bash) mastery is essential for day-to-day operational tasks and automation.
Monitoring: Extensive experience implementing and maintaining modern monitoring systems such as Prometheus, Victoria Metrics, and Riemann.
Logging: Proficiency with centralised log management using Loki for log ingestion, enrichment, lifecycle management, and providing a search/view platform.
Visualisation: Expertise in creating and managing dashboards for visualisation and alerting using Grafana.
IaC: Mastery of Terraform for writing new component configurations and building automation for BAU (Business As Usual) tasks.
Configuration Management: Strong experience with configuration management tools like SaltStack (or Ansible) for automated deployment and configuration of services on VMs.
High-Availability Data Stores: Hands-on experience setting up, managing, and scaling high-availability databases like Mysql and Aerospike.
Time-Series/Search: Familiarity with Elastic Search and time-series databases like InfluxDB.
Replication/DR: Expertise in database replication between different regions, managing database migrations, setting up circular replication, and ensuring data sync during system and network issues.
Web/Proxy: Expert management of critical infrastructure components like Nginx and HA Proxy, including proxy management, endpoint addition, header configuration, and writing rewrite rules.
Messaging/Container: Experience with messaging queues like RMQ (RabbitMQ) and containerization technology like Docker.
Networking Services: Deep knowledge of DNS and other core network protocols.
Ownership and Accountability: A proactive approach to identifying and solving infrastructure challenges before they impact service availability.
Communication: Excellent written and verbal skills for documenting procedures, creating runbooks, and communicating with technical and non-technical stakeholders.
Mentorship: (For senior roles) Ability to mentor junior engineers and promote SRE best practices across the organisation.
SLO/SLA Management: Experience defining, monitoring, and meeting Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for critical services.
Toil Reduction: A commitment to measuring and actively reducing operational toil through automation (e. g., using SRE's Toil Reduction framework).
Cost Optimisation: Experience in identifying and implementing cloud resource optimisation and cost-saving measures within the Azure environment.

Experience
7-11 yrs

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App