Live opening · Posted 27 days ago

Site Reliability Engineer - 3 (Big Data)

PhonePe · Bangalore
Instahyre 7-11 yrs
You are 27 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 27 days ago
CompanyPhonePe
LocationBangalore
Experience7-11 yrs
SourceInstahyre
Listed27 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
0 min from Instahyre publishing this role to us finding it
1 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
15,908 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

Responsibilities:
Manage, maintain, and support incremental changes to Linux/Unix environments.
Lead on-call rotations and incident responses, conducting root cause analysis and driving postmortem processes.
Design and implement automation systems for managing big data infrastructure, including provisioning, scaling, upgrades, and patching clusters.
Troubleshoot and resolve complex production issues while identifying root causes and implementing mitigating strategies.
Design and review scalable and reliable system architectures.
Collaborate with teams to optimize overall system performance.
Enforce security standards across systems and infrastructure.
Set technical direction, drive standardization, and operate independently.
Ensure availability, performance, and scalability of systems and services through proactive monitoring, maintenance, and capacity planning.
Resolve, analyze, and respond to system outages and disruptions and implement measures to prevent similar incidents from recurring.
Develop tools and scripts to automate operational processes, reducing manual workload, increasing efficiency, and improving system resilience.
Monitor and optimize system performance and resource usage, identify and address bottlenecks, and implement best practices for performance tuning.
Collaborate with development teams to integrate best practices for reliability, scalability, and performance into the software development lifecycle.
Stay informed of industry technology trends and innovations, and actively contribute to the organization's technology communities.
Develop and enforce SRE best practices and principles.
Align across functional teams on priorities and deliverables.
Drive automation to enhance operational efficiency.
Requirements:
Over 7 years of experience managing and maintaining distributed big data ecosystems.
Strong expertise in Linux, including IP, iptables, and IPsec.
Proficiency in scripting/programming with languages like Perl, Golang, or Python.
Hands-on experience with the Hadoop stack (HDFS, HBase, Airflow, YARN, Ranger, Kafka, Pinot).
Familiarity with open-source configuration management and deployment tools such as Puppet, Salt, Chef, or Ansible.
Solid understanding of networking, open-source technologies, and related tools.
Excellent communication and collaboration skills.
DevOps tools: SaltStack, Ansible, Docker, Git.
SRE logging and monitoring tools: ELK stack, Grafana, Prometheus, OpenTSDB, OpenTelemetry.
Good to Have:
Experience managing infrastructure on public cloud platforms (AWS, Azure, GCP).
Experience in designing and reviewing system architectures for scalability and reliability.
Experience with observability tools to visualize and alert on system performance.

Experience
7-11 yrs

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App