Live opening · Posted 6 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Lead Monitoring & Observability Engineer (IC3)
Role Summary
The Lead Monitoring & Observability Engineer is a senior technical contributor responsible for monitoring reliability, telemetry quality, automation, and operational visibility across enterprise infrastructure and application environments.
This role combines hands-on engineering with technical leadership across observability, network monitoring, automation, and SRE practices. The engineer will establish monitoring and performance standards, improve alert quality and service reliability, build scalable observability solutions, and partner with product, infrastructure, networking, cloud, security, and engineering teams.
The ideal candidate is highly technical, hands-on, and comfortable operating as a monitoring/observability SME while mentoring engineers and influencing technical direction.
Key Responsibilities
Monitoring & Observability
Establish and maintain SLOs, availability targets, performance objectives, reliability metrics, and monitoring coverage standards.
Design and maintain actionable dashboards, alerts, logs, metrics, traces, telemetry pipelines, and service-health reporting.
Monitor platform and application health and implement improvements to reliability, performance, alert quality, and operational stability.
Develop reusable observability patterns, reference architectures, onboarding standards, and validation practices.
Evaluate monitoring and observability tools, integrations, and emerging technologies.
Hands-On Engineering & Automation
Design, configure, and maintain monitoring and observability platforms across enterprise environments.
Develop infrastructure-as-code and automation using Terraform, Ansible, scripting, monitoring-as-code, and CI/CD pipelines.
Build scalable monitoring, alerting, logging, metrics, tracing, and telemetry solutions.
Automate monitoring onboarding, dashboard creation, alert configuration, telemetry collection, and operational workflows.
Reduce manual operational processes through automation and self-service tooling.
Network Monitoring
Define monitoring standards for network availability, reachability, latency, packet loss, interface health, capacity, routing, and service dependencies.
Build and maintain network-monitoring dashboards, alerts, service-health views, and operational workflows.
Monitor critical network devices, services, interfaces, and customer-impacting dependencies.
Establish validation standards for network-device onboarding, telemetry collection, alert quality, dashboard completeness, and operational readiness.
Support migration from OpsRamp to a new network-monitoring platform, including requirements, technical design, migration planning, testing, validation, and operational handoff.
Evaluate network-monitoring tools, integrations, automation opportunities, and emerging observability requirements.
Correlate network telemetry with application, infrastructure, and customer-impact signals to improve incident triage and resolution.
Incident Response & Operational Excellence
Improve alert signal quality, reduce alert noise, and establish clear ownership, escalation paths, and runbook guidance.
Participate in and lead high-severity incidents involving monitoring and observability platforms.
Perform root-cause analysis using telemetry, infrastructure conditions, application behavior, and operational events.
Improve monitoring capabilities based on evolving product and engineering requirements.
Support service owners in prioritizing observability improvements for critical applications and infrastructure.
Technical Leadership & Mentorship
Partner with AppDev, Security, Networking, Cloud, SRE, and infrastructure teams to deliver shared monitoring and observability solutions.
Mentor junior engineers on engineering best practices, monitoring fundamentals, TDD, Agile practices, and operational workflows.
Provide code reviews, architecture guidance, design feedback, and technical enablement.
Drive adoption of observability standards and reusable implementation patterns.
Provide technical leadership across monitoring architecture, telemetry quality, alerting, and operational readiness.
Required Qualifications
Strong experience in monitoring and observability engineering across enterprise infrastructure and application environments.
Hands-on experience with Elastic/ELK, dashboards, alerting, logging, metrics, tracing, telemetry pipelines, cloud infrastructure, networking, or similar technologies.
Experience defining availability and performance objectives and implementing monitoring-driven improvements.
Strong automation experience using Terraform, Ansible, scripting, monitoring-as-code, or cloud-platform automation.
Experience designing, implementing, and maintaining monitoring, alerting, dashboards, and service-health reporting.
Strong understanding of observability across logs, metrics, traces, telemetry, alerting, and incident response.
Experience with enterprise network monitoring, network operations, or network observability.
Working knowledge of TCP/IP, DNS, HTTP/S, SNMP, ICMP, routing, switching, firewalls, load balancing, VPN, and WAN connectivity.
Experience monitoring network devices, interfaces, availability, latency, packet loss, bandwidth, capacity, and reachability.
Experience designing network dashboards, alerts, escalation workflows, and service-health views.
Strong troubleshooting, analytical, and root-cause-analysis skills.
Ability to evaluate technical alternatives and influence architecture and platform decisions.
Experience working in Agile environments and using TDD or similar development practices.
Strong communication, collaboration, and mentoring skills.
Preferred Qualifications
Experience with distributed tracing, OpenTelemetry, or modern observability platforms.
Experience with DevOps/SRE practices, including SLOs, runbooks, deployments, incident response, and post-incident improvements.
Experience with multi-cloud or hybrid-cloud environments.
Experience with OpsRamp, SolarWinds, Elastic/ELK, Grafana, ServiceNow, or similar monitoring and incident-management platforms.
Experience with synthetic monitoring, customer-journey monitoring, or service-level reporting.
Experience leading an enterprise network-monitoring migration, platform consolidation, tool evaluation, or proof of concept.
#Remote
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.