Live opening · Posted 2 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Responsibilities:
Design and develop a next-generation scalable observability platform for modern cloud-native and hybrid infrastructures that works in tandem with AI agents.
Create intelligent AI agents to analyse logs, traces, and metrics in real time, delivering automated insights and remediation.
Build scalable and fault-tolerant AI agent frameworks.
Engineer and optimise large-scale analytics pipelines to process high-velocity telemetry data.
Build resilient distributed systems with high reliability, performance, and fault tolerance.
Implement and fine-tune LLMs for natural language querying and automated troubleshooting.
Partner with ML engineers to streamline AI model deployment and management.
Requirements:
Strong programming skills in Python and Golang (experience with Rust is a plus).
Track record of building distributed systems and large-scale analytics pipelines.
Hands-on experience with cloud infrastructure (AWS, GCP, or Azure) and Kubernetes.
Deep understanding of observability technologies (Prometheus, OpenTelemetry, Grafana, Elastic, etc. )
Knowledge of LLMs, AI agents, and agent frameworks like LangChain; AutoGen is a plus.
Experience with stream processing and real-time data processing frameworks.
Proficiency in database technologies (SQL and NoSQL, ClickHouse, Time-Series DBs).
5+ years of relevant experience in building highly scalable systems.
Bachelor's degree in Computer Science, Engineering, or a related field (Master's/PhD is a plus).
Experience
5-9 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.