Live opening · Posted 2 days ago

Site Reliability Engineering II

Bright Money · Bangalore
Instahyre 4-6 yrs
You are 2 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 2 days ago
CompanyBright Money
LocationBangalore
Experience4-6 yrs
SourceInstahyre
Listed2 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
14 min from Instahyre publishing this role to us finding it
1 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
16,383 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

The Infrastructure/SRE team is responsible for building, managing, and scaling Bright Money's cloud infrastructure to ensure our production systems are reliable, secure, scalable, and cost-efficient. The role involves owning AWS infrastructure, improving CI/CD and automation, strengthening observability and security, and ensuring the platform is prepared for incidents and disaster recovery.
You will work closely with engineering teams to operate production systems, automate infrastructure processes, improve system reliability, manage cloud costs, and build robust disaster recovery and failover capabilities.
Responsibilities:
Infrastructure Operations: Own and operate core production infrastructure on cloud platforms, ensuring high availability, observability, and scalability.
Incident Management: Lead incident response and author formal Root Cause Analysis (RCA) reports. Support the rollout of our new incident management platform and automated runbooks.
CI/CD and Automation: Design and optimise robust CI/CD pipelines. A major H2 goal is standardising pipelines across all services following our Python upgrade and containerization tracks.
Security and Compliance: Implement infrastructure security compliance, including IAM roles, SCPs, and our upcoming Identity Platform (Teleport) rollout.
Observability: Maintain monitoring dashboards and alerting. Support the revamp of our VictoriaMetrics HA stack and ELK log optimisation.
FinOps: Lead cloud cost analysis and resource tagging tracks to maintain efficient architecture.
Disaster Recovery: Lead the build-out of cross-region replicas and failover procedures to meet agreed RPO/RTO targets per service tier.
Requirements:
Cloud Expertise: Strong experience with AWS (required) and Infrastructure-as-Code tools like Terraform.
Scripting: Proficiency in Python and Bash. Experience with the Django framework is a plus to support our internal tooling.
Observability Stack: Deep knowledge of Prometheus, Grafana, VictoriaMetrics, and ELK/OpenSearch.
Systems Knowledge: Strong analytical skills across database (RDS) and message queue (RabbitMQ) systems.
Security Mindset: Familiarity with SSO/IdP integration, secret management (Vault/AWS Secrets Manager), and vulnerability remediation.
Containerization: Hands-on experience with Kubernetes (EKS) and orchestrating multi-environment clusters. This is critical as we move toward full containerization of workloads in H2

Experience
4-6 yrs

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App