Live opening · Posted 11 days ago

Sr. Manager, Data Reliability Engineering (8+ years)

Visa · IN - Bengaluru, India
Workday
You are 11 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 11 days ago
CompanyVisa
LocationIN - Bengaluru, India
SourceWorkday
Listed11 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
1 min from Workday publishing this role to us finding it
26 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
19,707 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

About Us
Visa is a world leader in payments technology, facilitating transactions between consumers, merchants, financial institutions and government entities across more than 200 countries and territories, dedicated to uplifting everyone, everywhere by being the best way to pay and be paid.
At Visa, you'll have the opportunity to create impact at scale — tackling meaningful challenges, growing your skills and seeing your contributions impact lives around the world.
Join Visa and do work that matters – to you, to your community, and to the world. Progress starts with you.
Job Description
Job Summary
The Senior Manager, Site Reliability Engineering (SRE) – Data Platforms leads teams responsible for ensuring the reliability, scalability, security, performance, and operational excellence of data platforms, data infrastructure, and data-intensive production systems. This role supports the Data area by partnering with data engineering, analytics, machine learning, product, risk, compliance, security, and infrastructure teams to operate and improve the systems that capture, process, store, move, and serve critical business data.
This role oversees the development and execution of reliability strategies, infrastructure automation, observability practices, incident management processes, and operational standards for data pipelines, data platforms, distributed processing systems, cloud data services, streaming platforms, orchestration tools, and data storage environments. The Senior Manager ensures that data services meet expectations for availability, latency, throughput, recoverability, security, compliance, and data integrity.
This role is responsible for establishing processes, automation, monitoring, alerting, service-level objectives, incident response practices, disaster recovery capabilities, capacity planning, and operational excellence frameworks based on business and technical requirements. The Senior Manager guides teams in improving platform reliability, reducing toil, troubleshooting complex production issues, and supporting resilient data operations across high-volume, distributed environments.
The Senior Manager acts as a leader for SRE, infrastructure, platform, or operations engineering teams supporting the Data organization by removing blockers, enabling productivity, improving operational maturity, and fostering an inclusive, high-performance engineering culture. The role also involves coaching and developing engineers to adapt to evolving cloud platforms, data technologies, automation tools, observability frameworks, reliability practices, and AI-driven workflows. This leader partners closely with software engineering, data engineering, product management, QA, risk, compliance, security, infrastructure, and business stakeholders to deliver safe, resilient, transparent, and customer-centric data solutions.
All roles require digital fluency, including the ability to work with emerging technologies such as Generative AI tools — for example, ChatGPT and Claude Cod — to support everyday work, improve operational efficiency, accelerate documentation, assist with troubleshooting, and enhance engineering productivity.
Key Responsibilities
Lead the design, development, deployment, and operation of scalable, reliable, secure, and highly available SRE solutions supporting data platforms, data pipelines, and data infrastructure.
Oversee the creation and maintenance of automation, monitoring, alerting, observability, incident response, disaster recovery, and reliability engineering practices for data systems across multiple platforms and services.
Guide teams in troubleshooting and resolving complex production issues involving data pipelines, distributed processing systems, storage platforms, orchestration tools, streaming systems, infrastructure services, performance bottlenecks, and service degradations.
Drive the adoption of automation, CI/CD, infrastructure as code, quality engineering, observability, reliability testing, and operational excellence practices across the Data area.
Coach and mentor SRE, infrastructure, platform, or operations engineering teams supporting Data, fostering a culture of continuous improvement, accountability, learning, reliability, data integrity, and technical excellence.
Collaborate with data engineering, analytics, machine learning, product, security, risk, compliance, QA, infrastructure, and business partners to ensure solutions meet business priorities, customer expectations, regulatory requirements, security standards, data governance expectations, and operational risk controls.
Ensure adherence to reliability, resilience, security, compliance, access management, disaster recovery, change management, and data protection standards throughout the technology and data lifecycle.
Lead the implementation of service-level indicators, service-level objectives, error budgets, capacity planning, performance optimization, incident metrics, reliability dashboards, and operational health checks for data platforms and data services.
Partner with data engineering and platform teams to improve the reliability of batch pipelines, real-time streaming workloads, machine learning pipelines, data ingestion frameworks, transformation jobs, metadata processes, and reporting or analytics services.
Oversee the development and maintenance of technical documentation, runbooks, architecture diagrams, operational procedures, post-incident reviews, production-readiness checklists, and best practices for code quality, change review, and data platform supportability.
This is a hybrid position. Expectation of days in the office will be confirmed by your Hiring Manager.
Qualifications
Basic Qualifications
8+ years of relevant work experience and a Bachelor's degree, OR 11+ years of relevant work experience.
Experience in leading teams in the design, development, and deployment of large-scale engineering solutions.
Experience with cloud infrastructure, distributed systems, observability platforms, and reliability engineering practices.
Experience in building and maintaining scalable infrastructure, automation frameworks, incident response processes, and service reliability programs.
Experience with programming, scripting, and infrastructure tools such as Python, Go, Bash, Terraform, Kubernetes, CI/CD platforms, and version control systems such as Git.
Experience troubleshooting and resolving issues in complex, high-volume, highly available production environments.
Experience implementing security, compliance, access control, disaster recovery, and operational risk management standards.
Experience coaching, mentoring, and developing SRE, infrastructure, platform, or operations engineering teams.
Experience collaborating with engineering, product, security, infrastructure, and business stakeholders to improve reliability, scalability, and operational excellence.
Experience developing and maintaining technical documentation, runbooks, post-incident reviews, service-level objectives, operational procedures, and reliability dashboards.
Preferred Qualifications
9 or more years of relevant work experience with a Bachelor’s Degree, or 7 or more years of relevant experience with an Advanced Degree, or 3 or more years of experience with a PhD.
Experience supporting large-scale data platforms and data infrastructure in production environments.
Experience with cloud platforms and services, such as AWS, Azure, or GCP, used to operate reliable and scalable data systems.
Experience with Site Reliability Engineering practices, including monitoring, alerting, incident management, service-level objectives, error budgets, and post-incident reviews.
Experience improving the reliability of data pipelines, distributed processing systems, streaming platforms, orchestration tools, storage platforms, and analytics environments.
Experience troubleshooting complex production issues across data platforms, infrastructure, applications, and distributed systems.
Experience with automation, CI/CD, infrastructure as code, observability tools, containerization, and orchestration platforms.
Experience supporting systems requiring high availability, low latency, scalability, resiliency, and disaster recovery.
Experience partnering with data engineering, analytics, machine learning, product, security, risk, compliance, QA, and infrastructure teams.
Experience leading cross-functional SRE, platform, infr

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App