Live opening · Posted 1 day ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Responsibilities:
Design and build applications to improve the resiliency of the critical systems.
Work across multiple technologies and applications.
Collaborate with a talented team in a fast-paced environment, learning and helping others learn.
Proactively engage application owners and drive conversations to unblock delivery.
Design and implement observability solutions to build monitoring dashboards, alerting, and health-check mechanisms to provide real-time visibility into failover readiness and execution.
Recommend and establish best practices, evaluate current processes, identify gaps, and propose improvements for failover patterns, automation standards, and operational runbooks.
Document everything and create clear, comprehensive technical documentation, architecture diagrams, runbooks, and onboarding guides that enable team scalability and knowledge transfer.
Requirements:
3+ years of hands-on software engineering experience across multiple technologies, languages, and system layers.
Strong first-principles understanding of distributed systems, fault tolerance, and failure modesnot just framework familiarity, but genuine depth.
GitLab CI/CD expertise.
Python/Bash scripting; strong YAML skills.
AWS and Kubernetes experience.
Familiarity with secret management (CyberArk, Vault).
Accountability mindset: you own problems end-to-end, you don't wait to be unblocked, and you escalate with context and a proposed path forward.
Strong documentation skills: ability to translate complex systems into clear, actionable guides.
Self-driven: You take ownership, find answers yourself, and don't wait to be told what to do next.
First-principles thinker: When something breaks in an unfamiliar system, you reason from fundamentals. You don't just apply patterns; you understand why the pattern exists.
Fast learner: You ramp quickly on new tools and ecosystems with minimal guidance.
Independent operator: You can engage app teams directly, extract what you need, and fill gaps through your own research.
Fast, iterative, and comfortable with ambiguity: You ship something workable quickly, learn from it, and improve. You don't need the perfect spec to start
Relationship builder: You build trust with stakeholders and drive conversations forward.
Communicator with standards: You write clearly, document proactively, and treat your teammates' time as valuable.
Continuous improver: You don't just execute; you identify what's suboptimal and propose better ways of doing things, then follow through.
Knowledge sharer: You believe documentation is a first-class deliverable, not an afterthought.
Nice-to-Have:
Ansible and failover experience.
Telecom or large enterprise environment experience.
Experience with observability platforms (Splunk, Grafana, Prometheus, OTEL).
Experience using AI coding tools (Claude, GitHub Copilot, ChatGPT) as a genuine productivity multiplier, not just having tried them but having integrated them into your workflow.
Experience
3-6 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.