Live opening · Posted 7 hours ago

Senior Lead Site Reliability Engineer - Cloud Data & Databricks

JPMorgan Chase · Hyderabad, Telangana, India
Oracle
You are 7 hours behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 7 hours ago
CompanyJPMorgan Chase
LocationHyderabad, Telangana, India
SourceOracle
Listed7 hours ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
0 min from Oracle publishing this role to us finding it
12 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
15,854 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

Join us to shape the future of data and analytics, leveraging your expertise to deliver impactful technology solutions. Experience career growth and make a difference in a collaborative environment.
As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the AIML Data Platforms and Chief Data and Analytics Team, you will develop and deliver advanced technology products focused on data and analytics. You will tackle complex cloud data platform challenges, especially around Data Lake Tools, and work in an agile environment, collaborating with cross-functional teams. You will help drive the firm’s data and analytics journey, ensuring quality, integrity, and security of data, and leveraging AI/ML technologies to support commercial goals. You will contribute to a culture of innovation and operational excellence.
Job responsibilities
Lead a team of SREs to design, implement, and maintain a managed AWS Databricks platform, providing engineering and operational support to Application/Engineering teams
Perform platform design, set-up and configuration, workspace administration, and resource monitoring
Create multi-AZ, multi-region, and multi-cloud resiliency strategies for business-critical products and services
Lead evaluation sessions with external vendors, startups, and internal teams to assess architectural designs and technical credentials
Drive continuous improvement in system observability, alerting, and capacity planning
Collaborate with engineering and data teams to optimize infrastructure and deployment processes, focusing on automation and operational excellence
Execute creative software solutions, design, development, and technical troubleshooting
Uses enterprise-authorized AI capabilities within the work environment to accelerate reliability design and operational decisioning (e.g., incident/post-incident analysis and requirements traceability), validating outputs and handling operational data according to sensitivity and security requirements.
Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., testing/validation automation and production readiness), ensuring traceability/auditability, resiliency, and security controls.
Develop secure, high-quality production code, and review/debug code written by others to ensure reliability and correctness.
Apply SRE best practices to improve reliability, scalability, and performance; eliminate or automate recurring issues; and maintain incident response procedures including root cause analysis and postmortems.
Required qualifications, capabilities and skills
Formal training or certification on software engineering concepts and 10+ years applied experience
Strong understanding of SRE principles, including SLIs, SLOs, error budgets, and incident management
Experience with monitoring tools, automation frameworks, and CI/CD pipelines
Proficient in Python application program development with use of automated unit testing
Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve reliability engineering workflows with strong validation habits and awareness of data sensitivity.
Ability to set team practices for safe AI usage in operations (e.g., review/approval expectations and escalation paths) while maintaining resiliency, security, and auditability outcomes.
Experience with Terraform development and understanding of Terraform enterprise
Experience in delivering system design, application development, testing, and operational stability
Knowledge of Big Data distributed compute frameworks like Spark, Glue, MapReduce
Excellent troubleshooting, analytical, and communication skills
Preferred qualifications, capabilities and skills
Experience in Data pipelines using Spark
Exposure to AWS & Databricks Platform administration
Knowledge of containerization (Docker, Kubernetes) and orchestration
Familiarity with distributed systems and large-scale data processing

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App