Live opening · Posted 12 hours ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Roles & Responsibilities
Ensure production reliability and performance by monitoring service availability, latency, capacity, and overall system health while managing SLOs, SLIs, and error budgets.
Automate operational processes and reduce toil by developing tools and scripts for deployment, infrastructure management, incident response, capacity planning, and other repetitive operational tasks.
Manage incident response and root-cause analysis, participate in on-call rotations, troubleshoot production issues, and conduct blameless postmortems to implement long-term corrective actions.
Design and improve scalable infrastructure by collaborating with software engineering teams on system architecture, distributed systems, reliability, scalability, and performance requirements.
Manage cloud infrastructure and deployment processes using technologies such as Kubernetes, Terraform, Ansible, and CI/CD pipelines, including canary releases and automated deployment practices.
Implement observability, capacity planning, and performance optimization across production systems, using monitoring and logging platforms such as Datadog or Splunk and optimizing SQL/NoSQL databases and distributed services.
Qualifications
Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related technical field, or equivalent practical experience.
4+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Cloud Infrastructure, or a closely related role, with strong exposure to production environments.
Strong programming or scripting skills in at least one language such as Python, Go, Java, or C++, with the ability to develop automation tools and solve complex technical problems.
Strong knowledge of Linux/Unix systems, networking, and distributed systems, including TCP/IP, DNS, load balancing, service reliability, and troubleshooting across distributed environments.
Hands-on experience with cloud-native infrastructure and Infrastructure as Code, including Kubernetes, Terraform, Ansible, CI/CD, and automated deployment practices.
Experience with observability and database technologies, including tools such as Datadog/Splunk and both relational and NoSQL databases, along with strong analytical, debugging, incident-management, and performance-tuning skills.
Skills: ci,infrastructure,cloud,reliability,distributed systems
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.