Live opening · Posted 27 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
The role is for an Autonomous Systems Engineer (Platform) focused on improving the reliability, operability, and evolution of the internal engineering platform. It combines platform engineering, site reliability, and intelligent automation, with a strong emphasis on reducing manual work, improving observability, and enabling safe AI-driven or agent-based automation at scale.
Responsibilities:
Closed-Loop Resilience: Enable "detect, diagnose, remediate, validate" automation systems to improve system resilience and ensure customer-impacting issues are addressed proactively.
AI-Driven Automation: Build and operate platform automation and AI-powered (agent-based) workflows to minimize manual operational effort and move toward a self-managing infrastructure.
Reliability and Incident Leadership: Own the performance, availability, and reliability of the internal platform and critical services, which includes participating in on-call rotations and leading incident triage, debugging, and root cause analysis.
Safety and Guardrail Design: Design and implement validation pipelines, safety mechanisms, and guardrails specifically for automated and AI-generated changes to infrastructure and code to ensure production stability.
Requirements:
Production Operations: At least 5+ years of experience operating large-scale distributed systems or production platforms.
Programming and Scripting: Proficiency in Java, Python, Go, or Shell.
Incident Management: A proven track record in production debugging, on-call operations, and leading root cause analysis.
System Architecture: A solid understanding of distributed systems architecture and common failure modes.
Linux Expertise: Practical, working knowledge of Linux-based production environments.
AI/Automation Integration: Experience building or integrating AI-driven (agent-based) automation frameworks.
Essential Tooling and Frameworks:
Observability Tools: Hands-on experience with tools for monitoring, alerting, logging, and tracing.
CI/CD Pipelines: Expertise in building and maintaining automated release processes and deployment pipelines.
Lifecycle Management: Familiarity with managing platform upgrades and dependency lifecycles.
Safety and Governance: Skills in designing safety and validation layers for automated systems.
Nice-to-haves: SRE best practices, chaos/resilience testing, self-healing automation, governance for automated systems, and large-scale platform standardization experience.
Expectations:
Reliability and production operations expertise: This role is centered on owning the reliability, availability, and performance of the platform, so strong experience with production systems, incident response, on-call operations, debugging, and root cause analysis is essential.
Strong automation mindset: A major expectation is reducing toil through automation and building AI-powered or agent-based workflows. That makes automation-first thinking one of the most important qualities for success in the role.
Distributed systems and platform engineering knowledge: The role requires a solid understanding of distributed systems architecture, failure modes, platform upgrades, dependency management, and lifecycle operations.
Observability and operational excellence: Experience with monitoring, alerting, logging, tracing, SLIs, and SLOs is critical because the role is expected to improve visibility and guide engineering decisions using reliability data.
Cross-team collaboration and communication: Because the role partners directly with engineering teams on service design, resilience, and operability, strong communication and collaboration skills are also very important.
Experience
5-9 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.