Live opening · Posted 6 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Cloud DevOps Engineer
Location: Chennai, UK
Work Mode : Hybrid
Job purpose
We are seeking a Site Reliability Engineer to safeguard the reliability, observability and resilience of critical platforms supporting a Global Capital Markets platform-modernization program. The role ensures that Front Office Trading systems and the emerging target-state services — spanning event-driven integration, data-lake / Databricks platforms, reporting, migration and the front-to-back trade lifecycle for FX Cash, Cleared IRS and future products — run reliably across the trading day and its critical batch cycles. The ideal candidate blends strong engineering with an operations mindset, automating toil, engineering for resilience, and improving supportability before services reach production.
Key responsibilities
Define and implement service-health metrics, monitoring, alerting and dashboards to provide end-to end observability.
Support incident response, problem management and operational readiness, including root-cause analysis and post-incident reviews.
Identify and remediate reliability, performance and resilience risks across critical trading and post trade services.
Monitor batch / EOD cycles and time-critical processes, detecting failures, delays and job / queue issues early.
Engineer automation to reduce operational toil (self-healing, runbooks, scheduling) and support capacity planning and performance tuning & and improve MTTR.
Partner with engineering teams to improve supportability, deployment safety and production readiness before release.
Contribute to SLIs / SLOs, error-budget management and continuous improvement of operational resilience and BCP.
Support Kubernetes-based platforms including monitoring cluster health, workload performance and platform reliability.
Key competencies
Bachelor’s / Master’s degree with 8+ years in SRE / production-engineering / DevOps roles
supporting business-critical, high-availability systems.
Strong observability skills — monitoring, alerting and dashboards using tools such as
Prometheus/Grafana, ELK, Splunk, Datadog, Dynatrace, AppDynamics or Cloud-native monitoring platforms.
Incident and problem management experience, including on-call, escalation and RCA; scheduling tools (e.g., Autosys) and alerting (e.g., PagerDuty).
Automation and scripting (Python / Bash / PowerShell), with CI/CD and infrastructure-as-code familiarity.
Cloud experience (AWS and/or Azure), strong Linux fundamentals, and understanding of resilience, performance and capacity engineering.
Strong analytical, communication and cross-team collaboration skills, with attention to detail.
Understanding of disaster recovery (DR), high availability (HA), failover mechanisms and resilience engineering practices.
Preferred qualifications
Experience defining and operating against SLIs / SLOs and error budgets.
Exposure to containers / orchestration (Docker / Kubernetes) and streaming / event-driven platforms (Kafka / MQ).
Prior experience in Capital Markets or financial services, including awareness of settlement and stress-test batch processing.
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.