Live opening · Posted 7 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
About the Role
We are looking for a Lead Reliability Engineer to own and evolve The Hartford's enterprise Software Delivery Framework (SDF) platform. You will be the technical authority on Jenkins (primary), Harness, Git / GitHub Enterprise, uDeploy, Nexus, and SonarQube — driving platform stability, service reliability, and cost-efficiency through strong operational engineering and governance. An AI-first mentality is a core expectation. Cross-team collaboration is a must — this role partners with engineering, security, observability, release management, and infrastructure teams to align standards and unblock delivery.
Key Responsibilities
Own the full SDF platform lifecycle: Jenkins, Harness, Git / GitHub Enterprise, uDeploy, Nexus, and SonarQube.
Ensure platform stability and availability across SDF tooling through proactive reliability engineering practices.
Define and enforce enterprise reliability standards, operational controls, and compliance guardrails.
Drive incident prevention and rapid recovery: observability, early risk detection, runbook maturity, and resilience testing.
Apply AI-first approaches for reliability diagnostics, intelligent alert triage, and auto-remediation opportunities.
Own end-to-end RCA for Sev1/Sev2 SDF incidents — from detection through corrective action and verified closure; publish stakeholder summaries.
Drive alert noise reduction across SDF tooling — enforce signal-to-noise standards and measurable alert quality gates.
Lead cost optimization initiatives across tooling and infrastructure without compromising reliability or developer experience.
Coach junior engineers — define escalation criteria, build triage playbooks, and close knowledge gaps.
Cross-team collaboration — partner with application engineering, observability, security, release management, and infrastructure teams to align standards and unblock delivery.
Lead end-user experience outcomes as the centralized SDF platform owner, ensuring every reliability and stability improvement delivers a simpler, faster, and more consistent experience across all SDF tools.
Build and maintain a self-onboarding framework so new teams can independently adopt SDF operational best practices.
Manage SDF infrastructure via Terraform and Ansible; automate provisioning, health checks, and self-healing.
Track and report platform reliability and efficiency metrics: availability, MTTR, incident recurrence, and cost savings.
Platform & Tools
Jenkins Harness Git / GitHub Enterprise uDeploy Nexus SonarQube Terraform Ansible Harness Python / Groovy / Bash REST APIs Rally
Requirements
Jenkins administration (required)
Harness CD pipelines
Git / GitHub Enterprise
uDeploy
Nexus Repository Manager
SonarQube
Platform reliability engineering and operational governance
Service stability, resilience, and observability practices
Cross-team collaboration (required)
End-to-end RCA ownership
Cost optimization and efficiency mindset
Coaching and mentoring junior engineers
Terraform / Ansible (IaC)
AI-first problem-solving approach
Regulated industry experience a plus
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.