Live opening · Posted 2 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Responsibilities:
Design, build, and ship LLM-powered and agentic product features that enhance the team's efforts and outcomes.
Build agentic AI systems that reason over context, invoke tools, take real actions, and recover gracefully from failure.
Work on integrating the existing AI tools, and should know major AI frameworks and libraries.
Own service reliability and operational governance by defining SLA's, managing error budgets, and reporting reliability (MTTD, MTTR) to leadership for prioritisation, risk decisions and planning.
Architect and continuously optimise the observability of the platform using Kibana/Elastic (ELF) along with other observability tools like Prometheus, Grafana (dashboards, metrics, alert lifecycle), improving detection quality, reducing noise/toil, and enabling faster triage and measurable uptime improvements.
Engineer advanced alerting and automation capabilities with Kibana alerting and anomaly detections, and integrating response workflows (routing, runbooks, remediation scripts) to standardise on-call execution and accelerate restoration of services.
Lead incident response for customer-impacting issues across teams, coordination, communications, service restoration, and blameless RCA, then corrective actions that prevent recurrence and reduce operational risk.
Design, automate and validate Disaster Recovery and failover for critical services/journeys (RTO/RPO alignment, DR Drills), ensuring resiliency under failure scenarios and improving recovery.
Consult and partner with application teams by providing production readiness inputs (Resiliency patterns, availability, performance/capacity considerations) and driving platform enhancements that improve stability while optimising infrastructure and observability spend.
The core requirements for the job include the following:
Primary Skills:
Observability - ELK (Elastic/Kibana), Prometheus, Grafana, PromQL.
Software and automation - Java, Vertx, Python/Shell/Bash, Rest-SOAP API, Docker containerization, Kubernetes, Kafka.
Reliability and DR engineering - Distributed architecture and distributed system fundamentals, micro services, and event-driven architecture.
Cross-team coordination, incident triage and resolution, leadership and stakeholder management.
Secondary Skills:
Lang-chain, Langraph, RAG, MCP.
Experience with working on LLM's and integrating with the existing applications.
Python - FastAPI.
Cache - Redis.
Experience
6-10 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.