Live opening · Posted 5 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Summary:
We are seeking a highly skilled Software Engineer with a strong background in managing production incidents, debugging distributed systems, and implementing robust monitoring solutions. The ideal candidate will have extensive experience in programming and a keen interest in utilizing AI for enhancing system reliability.
Responsibilities:
Drive production incidents as the commander or primary responder with 24x7 on-call availability.
Debug distributed systems by correlating metrics, logs, and traces, and work comfortably with SLIs, SLOs, and error budgets.
Design and maintain Datadog dashboards, monitors, APM, tracing, and log pipelines.
Instrument services with Open Telemetry and design custom metrics, considering cardinality and cost.
Engage actively with AI engineering assistants to apply AI in reliability work, such as alert triage, log summarization, and RCA drafting.
Conduct capacity planning, load and performance testing, and ensure peak-event readiness.
Requirements:
Minimum of 8 years of experience in software engineering roles.
Experience in payments, fintech, or other high-availability regulated domains.
Required Skills:
Proficiency in Datadog for dashboards, monitors, APM, tracing, and log pipelines.
Experience with Open Telemetry for instrumenting services and designing custom metrics.
Strong programming skills in Java, Spring Boot, Node.js, and Python.
Preferred Skills:
Practical use of AI engineering assistants.
#AditiConsulting
# 26-05930
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.