Live opening · Posted 1 day ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Celigo is looking for a Senior Site Reliability Engineer to join a team of highly talented individuals committed to offering the best quality services and products in the area of business cloud computing (SaaS).
The core responsibilities for the job include the following:
Reliability Engineering:
Define, measure, and enforce SLIs, SLOs, and SLAs across platform services.
Manage error budgets and drive trade-off discussions with engineering teams.
Identify and reduce toil through automation with measurable outcomes.
Conduct capacity planning, forecasting, and scaling strategies.
Implement chaos engineering practices (Chaos Monkey, Gremlin, Litmus, and fault injection).
Observability and Monitoring:
Build, maintain, and continuously improve observability infrastructure metrics, logging, tracing, and APM.
Set up dashboards, alerting, and log analysis using tools such as Splunk, Datadog, Prometheus, and Grafana.
Proactively identify and resolve performance bottlenecks and reliability risks before customer impact.
Incident Management:
Own and respond to production incidents; lead bridge calls during outages.
Lead blameless post-mortems with root cause analysis (5 Whys, Ishikawa).
Track and improve MTTR, MTTD, and MTBF.
Write and maintain runbooks for operational procedures.
Manage on-call rotations using PagerDuty, Opsgenie, or VictorOps.
Performance and Scale:
Conduct load testing using JMeter, k6 Locust, or Gatling.
Profile and resolve CPU, memory, and I/O bottlenecks.
Apply distributed systems principles: CAP theorem, consensus, and eventual consistency.
Implement caching strategies using CDNs, Redis, and Memcached.
Infrastructure and Automation:
Automate infrastructure provisioning and deployments using Terraform and GitOps (ArgoCD).
Manage containerized workloads on Kubernetes and Docker.
Design and maintain CI/CD pipelines using Jenkins / Groovy DSL.
Implement deployment strategies: blue-green, canary, and rolling.
Requirements:
Strong experience in Python, Groovy/Jenkins DSL, and Shell Scripting (Bash).
Solid hands-on experience with AWS services (Lambda, EC2 S3 DynamoDB, etc. ).
Familiarity with RESTful API design and microservices architecture.
Proficient with version control systems (e. g., Git) and CI/CD pipelines.
Hands-on experience with observability and monitoring tools (e. g., Splunk, Datadog, Prometheus, Grafana, or equivalent), including setting up dashboards, alerts, and log analysis.
Strong understanding of SLIs, SLOs, and error budgets with experience defining and tracking reliability metrics.
Experience with production incident management, on-call rotations, and leading post-mortem processes.
Understanding of containerization technologies like Docker is a plus.
Excellent problem-solving skills and an ability to quickly learn and adapt to new technologies.
5-10 years of hands-on experience in software development, preferably in a cloud-based or product-driven environment.
Bachelor's degree in Computer Science or a related field, or equivalent practical experience.
Experience
7-11 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.