Live opening · Posted 7 hours ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Senior Principal Observability Engineer-IT
Be a part of a team that’s ensuring Dell Technologies' product integrity and customer satisfaction. Our IT Software Engineer team turns business requirements into technology solutions by designing, coding and testing/debugging applications, as well as documenting procedures for use and constantly seeking quality improvements.
What you’ll achieve:
You will be responsible for leading the deployment, optimization, and operational excellence of enterprise observability platforms while driving OpenTelemetry (OTEL) implementation strategies across the organization. You will work closely with architects, platform engineering teams, and business stakeholders to build scalable, resilient observability solutions that improve system performance, reliability, and operational insights.
Join us to do the best work of your career and make a profound social impact as a Senior Principal on our Software Engineer-IT team.
Take the first step towards your dream career
Every Dell Technologies team member brings something unique to the table. Here’s what we are looking for with this role:
Essential Requirements
12+ years of experience in Platform Engineering, DevOps, Site Reliability Engineering (SRE), or related infrastructure-focused roles.
Strong hands-on experience administering container orchestration platforms and managing production-scale infrastructure environments.
Deep expertise in OpenTelemetry (OTEL), including instrumentation, collectors, telemetry pipelines, and observability best practices.
Proven experience with observability platforms, telemetry data modeling, incident response, and monitoring strategies for high-volume environments.
Strong automation and scripting skills using Python, Bash, or infrastructure-as-code technologies, with experience optimizing relational and analytical databases.
Desirable Requirements
Experience deploying and operating enterprise observability platforms such as Splunk, Dynatrace, MIMIR, LangSmith, or similar solutions, including expertise with time-series databases, query optimization, APM, and distributed tracing in microservices environments.
Familiarity with LLMOps and LLM application concepts, including tracing, prompt evaluation, dataset management, prompt management, and experience with telemetry-related technologies such as caching systems, message queues, and streaming components.
You will:
Lead the deployment, lifecycle management, and optimization of enterprise observability platforms such as Splunk, Dynatrace, MIMIR, and LangSmith, including system upgrades and configuration management.
Design and implement telemetry collection strategies, leveraging OpenTelemetry (OTEL) collectors, instrumentation, and pipelines for metrics, logs, traces, and profiles.
Own platform performance and reliability, including database tuning, observability data modeling, storage optimization, and incident troubleshooting across infrastructure and application layers.
Automate operational processes through infrastructure-as-code and GitOps practices while implementing monitoring and alerting solutions for platform health.
Mentor engineers and collaborate with cross-functional teams to drive scalability, security, and observability best practices across the enterprise.
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.