Live opening · Posted 2 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
The Infrastructure/SRE team is responsible for building, managing, and scaling Bright Money's cloud infrastructure to ensure our production systems are reliable, secure, scalable, and cost-efficient. The role involves owning AWS infrastructure, improving CI/CD and automation, strengthening observability and security, and ensuring the platform is prepared for incidents and disaster recovery.
You will work closely with engineering teams to operate production systems, automate infrastructure processes, improve system reliability, manage cloud costs, and build robust disaster recovery and failover capabilities.
Responsibilities:
Infrastructure Operations: Own and operate core production infrastructure on cloud platforms, ensuring high availability, observability, and scalability.
Incident Management: Lead incident response and author formal Root Cause Analysis (RCA) reports. Support the rollout of our new incident management platform and automated runbooks.
CI/CD and Automation: Design and optimise robust CI/CD pipelines. A major H2 goal is standardising pipelines across all services following our Python upgrade and containerization tracks.
Security and Compliance: Implement infrastructure security compliance, including IAM roles, SCPs, and our upcoming Identity Platform (Teleport) rollout.
Observability: Maintain monitoring dashboards and alerting. Support the revamp of our VictoriaMetrics HA stack and ELK log optimisation.
FinOps: Lead cloud cost analysis and resource tagging tracks to maintain efficient architecture.
Disaster Recovery: Lead the build-out of cross-region replicas and failover procedures to meet agreed RPO/RTO targets per service tier.
Requirements:
Cloud Expertise: Strong experience with AWS (required) and Infrastructure-as-Code tools like Terraform.
Scripting: Proficiency in Python and Bash. Experience with the Django framework is a plus to support our internal tooling.
Observability Stack: Deep knowledge of Prometheus, Grafana, VictoriaMetrics, and ELK/OpenSearch.
Systems Knowledge: Strong analytical skills across database (RDS) and message queue (RabbitMQ) systems.
Security Mindset: Familiarity with SSO/IdP integration, secret management (Vault/AWS Secrets Manager), and vulnerability remediation.
Containerization: Hands-on experience with Kubernetes (EKS) and orchestrating multi-environment clusters. This is critical as we move toward full containerization of workloads in H2
Experience
4-6 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.