Live opening · Posted 4 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
This is a Fully Remote Job
1. About Our Client:
This organization operates in the technology consulting and software development industry, providing cloud, AI, data, and enterprise solutions across the United States. It addresses challenges related to scalable and reliable technology infrastructure by delivering expert services that support businesses in these areas.
2. About the Opportunity:
The Observability Engineer role is focused on designing and managing comprehensive observability platforms that enhance engineering teams' confidence in system performance. This position is critical for ensuring the usability, quality, and operational efficiency of monitoring solutions that transform telemetry data into actionable insights for engineering and business stakeholders.
3. Responsibilities:
Design and operate observability platforms covering metrics, logs, traces, events, and synthetic monitoring.
Architect deployments of Prometheus, Thanos, Mimir, Grafana, Loki, Tempo, OpenTelemetry, and Datadog for scalability and high availability.
Develop standards for service instrumentation, including OpenTelemetry adoption, metric naming, label cardinality, and structured logging.
Define and enforce SLOs, SLIs, and error budgets; build dashboards and alerts to support them.
Create alerting strategies that reduce noise and integrate with on-call tools such as PagerDuty and Opsgenie.
Manage large-scale time-series and log storage balancing retention, performance, and cost.
Design distributed tracing pipelines to assist in diagnosing latency and reliability issues.
Develop self-service tools and libraries to promote adoption of observability standards.
Drive cost management and label cardinality discipline in observability infrastructure.
Lead improvements in incident response readiness through dashboards, alert hygiene, and post-incident analysis tooling.
Collaborate with SRE and platform teams to integrate observability with deployment pipelines and delivery workflows.
Evaluate and recommend observability tools and vendors based on cost, capability, and maturity.
Mentor teams on observability best practices, debugging, and SLO-driven operations.
Maintain documentation, onboarding materials, and runbooks for the observability platform.
4. Requirements:
Bachelor’s degree in Computer Science or related field.
10+ years experience in SRE, platform engineering, or observability roles.
Extensive hands-on experience with Prometheus, Grafana, and at least one commercial observability platform (Datadog, New Relic, or Splunk).
Strong knowledge of OpenTelemetry, distributed tracing, and structured logging.
Proficiency in at least one programming language such as Go, Python, or Java.
Experience managing high-cardinality, high-throughput metrics and log pipelines.
Solid understanding of SLOs, error budgets, and SRE principles.
Experience integrating observability with CI/CD and incident management tools.
Good knowledge of Linux internals, networking, and container platforms.
Excellent communication and collaboration skills.
5. Pay Range and Compensation Package:
Salary range: $100,000–$160,000 annually.
Equal Opportunity Statement:
Equal Opportunity Statement: Our client is an equal opportunity employer. They celebrate diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, or national origin.
Note:
TalentHop is a recruitment partner of this role. Please note that all employment decisions, including candidate assessment, interviews, hiring, compensation, and employment terms, are made exclusively by the hiring employer.
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.