Live opening · Posted 15 hours ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Who are we?
At Finastra, we’re a global leader in financial services software, dedicated to expanding access to financial services and shaping what’s next for the industry. Our technology powers mission‑critical solutions across Lending, Payments and Universal Banking, supporting over 7,000 customers, including 80% of the world’s top 50 banks, in more than 110 countries.
What Success Looks Like
The successful candidate will help move the organization further toward an SRE operating model by reducing manual operations, improving observability, and engineering automation for recurring operational activities.
The engineer should be proactive rather than waiting for every task to be assigned. We are looking for someone who identifies repetitive work, reliability risks, monitoring gaps, and opportunities for automation and then takes ownership of developing the solution.
During production incidents, the engineer should actively investigate, communicate findings, recommend actions, and remain engaged through resolution.
The role will primarily focus on SRE, automation, DevOps, monitoring, and reliability engineering, while maintaining enough infrastructure and Windows expertise to contribute to the broader operational responsibilities of the team
Role Summary
We are seeking a hands-on Senior Site Reliability Engineer to improve the reliability, availability, automation, and operational efficiency of business-critical platforms.
The role focuses on Site Reliability Engineering practices, with a strong emphasis on automation, proactive monitoring, incident ownership, disaster recovery, and continuous improvement.
The position supports production infrastructure across Azure, and on-premises environments. The engineer will work across multiple technologies and platforms, including Windows infrastructure, with a focus on reducing manual operational work and improving service reliability.
This is a hands-on individual contributor role.
Key Responsibilities:
Site Reliability Engineering & Automation
Design and develop automation to reduce repetitive manual operational tasks.
Develop automation using PowerShell, Python, Ansible, APIs, and other appropriate technologies.
Integrate infrastructure automation into enterprise CI/CD and DevOps pipelines.
Implement infrastructure as code using Terraform and related technologies.
Develop reusable automation rather than one-off scripts.
Identify reliability risks and opportunities for automation and proactively drive improvements.
Develop automated health checks, validation, remediation, and reporting.
Disaster Recovery & Resilience
Design and implement automation for disaster recovery activities across multiple products.
Automate failover, failback, infrastructure validation, application validation, and customer validation.
Reduce manual intervention during DR exercises and recovery events.
Develop reusable DR automation that can be adopted across multiple products.
Improve system availability, resiliency, and recoverability.
Monitoring & Observability
Improve monitoring, observability, and alerting using Grafana and other enterprise monitoring platforms.
Develop and automate monitoring plugins and health checks.
Automate monitoring configuration and deployment.
Improve proactive detection of infrastructure and application issues.
Reduce unnecessary alerts and improve actionable monitoring.
Develop automated remediation where appropriate.
Infrastructure & Cloud Engineering
Support production infrastructure across Azure, virtualized, and on-premises environments.
Automate infrastructure provisioning, configuration, maintenance, and validation.
Support high availability and resilient infrastructure architectures.
Automate patching, vulnerability remediation, security hardening, and compliance validation.
Integrate existing Ansible-based infrastructure and security automation into enterprise DevOps pipelines.
Collaborate with cloud, network, database, middleware, security, application, and infrastructure teams.
Windows & Production Operations
Provide hands-on support for Windows Server infrastructure as part of the broader production environment.
Troubleshoot Windows services, Active Directory, Group Policy, DNS, networking, certificates, authentication, and operating system issues.
Support Windows patching, security hardening, and vulnerability remediation.
Develop automation to reduce manual Windows administration.
Participate in production maintenance and disaster recovery activities
Incident Management & Reliability
Participate in Sev1/2/3 production incidents and take technical ownership through service restoration.
Troubleshoot complex infrastructure and application-related issues.
Perform root cause analysis and implement permanent engineering solutions for recurring problems.
Remain actively engaged during incidents and communicate technical findings and recommended actions.
Identify patterns from incidents and convert them into monitoring, automation, or reliability improvements.
Embed AI into project delivery to drive productivity, automate routine activities, strengthen decision quality through trusted data, and deliver continuous process improvement while applying appropriate human judgement
Required Qualifications
Experience :5+ years of experience in Site Reliability Engineering, DevOps, cloud engineering, infrastructure engineering, or systems engineering.
Strong hands-on scripting and automation experience using PowerShell, Python, Bash, or similar languages.
Experience with CI/CD and enterprise DevOps platforms such as Azure DevOps, Jenkins, GitHub Actions, GitLab CI, or equivalent.
Experience with Ansible or similar configuration-management technologies.
Experience with Terraform or other infrastructure-as-code technologies.
Exper
More openings worth a look
Recently tracked roles with full details and direct application links.