Live opening · Posted 28 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Job Responsibilities
Improve the reliability, scalability, performance, and operational efficiency of assigned OCI Compute services and components.
Investigate and resolve complex production incidents; contribute to mitigation, recovery, RCA, and follow-up actions.
Own and improve service-level KPIs, SLOs, dashboards, alerting, deployment validation, and operational procedures for assigned systems.
Build automation and tooling to reduce operational toil and improve production safety.
Partner with development and infrastructure teams on service architecture, deployment, configuration, and reliability improvements.
Use observability, telemetry, event correlation, and AIOps capabilities to improve detection, diagnosis, and incident response.
Support upgrades, migrations, patching, capacity planning, performance tuning, security vulnerability management and production rollouts.
Troubleshoot distributed-system issues by analyzing service topology, dependencies, configuration, and failure modes.
Contribute to incident-management practices, operational readiness, and service ownership improvements.
Share technical knowledge and support team members through documentation, reviews, and collaboration.
Participate in a 12x7 on-call rotation and support response to customer-impacting incidents.
Career Level - IC3
Mandatory Skills
4–8 years of experience in SRE, Production Engineering, Cloud Operations, Systems Engineering, or a similar role.
Experience operating and improving highly available production systems.
Strong programming or scripting skills in Python, Java, Go, or similar languages.
Hands-on experience with Linux, cloud infrastructure, networking, compute, and storage.
Experience with production monitoring, alerting, dashboards, logs, metrics, and tracing.
Experience owning or improving service SLIs, SLOs, KPIs, and operational procedures.
Strong incident troubleshooting, RCA, debugging, and problem-solving skills.
Experience with deployment pipelines, release validation, automation, and change-management practices.
Understanding of distributed systems, service dependencies, capacity planning, and performance tuning.
Ability to work independently on technical problems and collaborate effectively with engineering teams.
Strong written and verbal communication skills.
Preferred Skills
Experience with OCI and cloud infrastructure services.
Experience with AIOps, anomaly detection, event correlation, predictive alerting, or automated remediation.
Experience with Kubernetes, containers, infrastructure-as-code, and CI/CD.
Experience with service migrations, fleet maintenance, upgrades, patching, or production rollouts.
Expe
More openings worth a look
Recently tracked roles with full details and direct application links.