Live opening · Posted 2 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
We continue to expand our highly talented Infrastructure teams and are seeking a candidate who has large-scale production experience with metrics, events, logs, and traces. In this role, you'll work with the Observability team to design, develop, and maintain solutions for R& D teams to monitor the behaviour and performance of their workloads, reduce the likelihood and impact of incidents, and better troubleshoot issues. Candidates with an operations (DevOps/SysOps/TechOps) background who have supported infrastructure at scale are encouraged to apply for this role. If you are a firm believer in Infrastructure as Code, continuous deployment/delivery practices, and helping teams understand how their services behave in real-world scenarios, then you might be a great fit!
Responsibilities:
System design, configuration, integration, deployment, and operations of Observability systems and tools.
These systems include a collection of metrics/logs/events from many backend services deployed across multiple AWS accounts and regions, and consumed by multiple teams.
Working with engineering teams to enable them to support their services from development to production.
Ensure our Observability platform exceeds goals for availability, capacity, efficiency, scalability, and performance, as well as meeting our internal SLOs.
Build the next generation of observability, integrating with Istio.
Write libraries and APIs that provide a simple, unified interface to other developers when they use our monitoring, logging, and event processing systems.
Enhance the existing alerting capabilities with Slack, Jira, and PagerDuty.
Helping build a continuous deployment system guided by metrics and data.
Bring anomaly detection into the observability stack.
Requirements:
Minimum of five years of experience.
Strong with Python or Go.
Cloud of choice, preference for AWS - Lambda, CloudWatch, IAM, EC2 ECS, S3
Solid understanding of Kubernetes.
Prometheus, PromQL, Thanos, AlertManager, Grafana, etc.
Strong knowledge of standard monitoring protocols/frameworks - Prometheus/Influx line format, SNMP, JMX, etc.
Elastic Stack, syslog, CloudWatch Logs.
Comfortable working with git, Github, and common CI/CD approaches.
IAC tooling like CloudFormation or Terraform.
Experience
5-9 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.