Live opening · Posted 1 day ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Job Title: Production Site Reliability Engineer (SRE) – Digital Payments
Role Overview
We are seeking a highly technical and driven Production SRE Engineer to manage and monitor mission-critical payment platforms including UPI, IMPS, and Payment Hub systems & various Payment Applications.
The role focuses on ensuring high availability, low latency, and seamless transaction experience for customers. The incumbent will collaborate with cross-functional teams (Engineering, Business, Compliance) and external regulators (RBI, NPCI) to maintain resilient and scalable payment infrastructure.
Key Responsibilities
Production Support & Incident Management
Provide L2/L3 production support for UPI, IMPS, and Payment Hub platforms & various Payment Applications.
Diagnose, triage, and resolve transaction failures, timeouts, and API disruptions.
Lead and participate in Major Incident Management (MIM) calls and ensure timely stakeholder communication.
Manage incidents, service requests, and problem tickets via Jira, ServiceNow.
Provide regular updates to internal stakeholders and regulatory bodies (NPCI/RBI) during critical issues.
Reliability Engineering & RCA
Perform deep-dive Root Cause Analysis (RCA) for recurring payment and system issues.
Implement preventive and corrective measures to improve system stability.
Drive SRE best practices including error budgets, SLIs/SLOs, and system resilience.
Experience in managing DR Drills & Documentations.
Reviewing the SOPs & its relative documentations.
Monitoring, Observability & System Engineering
Monitor key performance indicators:
Transaction success rates
Latency and response times
Failure trends and retries
Build and maintain dashboards using:
ELK Stack, Grafana, Kibana, Splunk, Datadog, Prometheus
Establish proactive alerting and anomaly detection mechanisms.
Work closely with engineering teams to design and optimize:
High-throughput payment switches
Routing logic
Settlement and reconciliation systems
Understand and support UPI architecture, IMPS rails, and payment orchestration layers & various Payment Applications.
Trace end-to-end transaction lifecycle across distributed systems.
External Partner & Regulatory Coordination
Coordinate with NPCI, partner banks, and TPAPs during outages, reconciliation issues, or network disruptions.
Lead integrations and ensure seamless onboarding of ecosystem participants.
Ensure compliance with:
RBI guidelines and data localization mandates
NPCI operational and technical standards
Technical Skills & Expertise
Payments Domain Knowledge
Strong expertise in:
UPI architecture and flows
IMPS rails
Payment gateway / switch systems
Payment Hub orchestration & various Payment Applications.
Core Technical Skills
Advanced SQL proficiency (joins, aggregations, stored procedures)
Strong hands-on experience in:
Linux/UNIX systems administration
Shell scripting
Ability to:
Read , Write and interpret All types documentation (SOPs, workflows, etc.)
Understand database schemas
Analyse system architecture and latency
Monitoring & Observability Tools
Hands-on expertise with:
ELK Stack (Elasticsearch, Logstash, Kibana)
Grafana, Prometheus
Splunk, Datadog
DevOps & Cloud
Experience with:
CI/CD pipelines, Containerization (Docker, Kubernetes)
Cloud platforms:
AWS / GCP / Azure
Key Competencies
Strong problem-solving and analytical skills
High ownership in production environments
Ability to work under pressure in real-time systems
Strong stakeholder communication and coordination
Focus on reliability, scalability, and performance
Must have can do, takes initiative, Drives end to end deliverables.
Proactive, solution-oriented mindset with ownership to resolve production issues under pressure.
Ability to clearly articulate incidents, updates, and RCA to stakeholders, leadership, and regulators.
Works effectively with cross-functional teams (engineering, product, partners, regulators).
Structured thinking to diagnose complex system failures and drive long-term fixes.
Ability to stay calm and effective during high-severity incidents and critical outages.
Knowledge of PCI-DSS compliance, Financial data governance & security best practices
Quickly adapts to changing technologies, incidents, and regulatory requirements in a fast-evolving payments ecosystem.
Precision in analysing logs, transactions, and system behaviour to avoid critical errors in production.
Effectively manage multiple incidents, tasks, and escalations in a high-pressure environment.
Ability to handle expectations and coordinate with internal teams, partners, and regulators efficiently.
Takes ownership to make quick, informed decisions during outages or critical production incidents & communications to various Stake holders including Regulatory.
More openings worth a look
Recently tracked roles with full details and direct application links.