Live opening · Posted 6 days ago

Principal Core Infrastructure Engineer

Oracle · BENGALURU, KARNATAKA, India
Oracle
You are 6 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 6 days ago
CompanyOracle
LocationBENGALURU, KARNATAKA, India
SourceOracle
Listed6 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
0 min from Oracle publishing this role to us finding it
5 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
16,405 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

Provides technical leadership for a set of services within OCI’s network monitoring and observability platform, which provides visibility into the health, performance, and behavior of OCI network infrastructure.
Designs and evolves highly scalable, reliable, and efficient distributed services for collecting, processing, storing, and analyzing network telemetry. Owns service architecture and technical decisions across telemetry ingestion, stream processing, metrics, alerting, andnetwork health monitoring.
Ensures owned services operate reliably at cloud scale, meeting requirements for availability, data integrity, latency, scalability, and operational efficiency. Identifies performance and reliability bottlenecks and drives architectural and engineering improvements.
Serves as a technical expert for owned services, leads resolution of complex production issues, and works closely with network engineering, infrastructure, and dependent service teams. Mentors engineers and contributes to architecture and engineering standards across the broader Network Monitoring organization.
Career Level - IC4
Responsibilities
Network Observability & Service Architecture
Own the architecture and technical evolution of critical services within OCI’s network observability platform.
Design distributed services that collect, process, aggregate, store, query, and expose network telemetry at cloud scale.
Design scalable telemetry ingestion and processing pipelines for high-volume and high-cardinality data.
Develop capabilities for monitoring network health, topology, state, performance, and device behavior.
Design solutions that enable timely detection, diagnosis, and isolation of network failures and performance degradation.
Ensure telemetry quality, completeness, freshness, accuracy, and availability within owned services.
Scalability & Distributed Systems
Design horizontally scalable and elastic services capable of supporting continued OCI infrastructure and traffic growth.
Identify and resolve performance, throughput, latency, storage, and scalability bottlenecks.
Design high-throughput streaming and event-processing systems, including partitioning, buffering, aggregation, backpressure, and failure handling.
Implement resilient state management, replication, synchronization, and recovery mechanisms.
Make appropriate engineering trade-offs across consistency, availability, latency, durability, performance, and cost.
Reliability & Operational Excellence
Define and meet SLOs for availability, durability, latency, data freshness, and correctness of owned services.
Design fault-tolerant services that operate through infrastructure failures, network disruptions, dependency failures, and software upgrades.
Define KPIs, telemetry, dashboards, and alerts required to understand service health and identify operational risks.
Lead diagnosis and resolution of complex production issues involving owned services and their dependencies.
Lead root-cause investigations and implement corrective actions that prevent recurrence.
Maintain high standards for operational readiness, capacity planning, deployment, upgrades, rollback, and recovery.
Participate in operational support and serve as an escalation point for critical issues involving owned services.
Network Monitoring & Analytics
Build capabilities for monitoring large-scale Layer 2 and Layer 3 network infrastructure.
Develop mechanisms to identify changes in network state, topology, reachability, performance, and device health.
Correlate telemetry across network devices and infrastructure services to improve fault detection and localization.
Develop approaches for identifying abnormal network behavior, telemetry gaps, capacity risks, and infrastructure failures.
Partner with network engineering teams to translate network behavior and operational requirements into monitoring capabilities.
Improve alert quality and signal-to-noise ratio to reduce unnecessary operational load.
Software Engineering & Automation
Design and implement high-quality, maintainable, and performance-sensitive production software.
Drive engineering practices for testing, code quality, automation, and safe software delivery within owned services.
Automate infrastructure provisioning, configuration, deployment, patching, upgrades, and rollback.
Improve service efficiency through performance optimization, capacity management, and reduction of operational toil.
Ensure security, compliance, and vulnerability remediation requirements are incorporated into service design and operation.
Technical Leadership & Collaboration
Provide technical leadership for projects and initiatives involving owned services.
Make and document architectural decisions and influence related designs across dependent services.
Lead complex technical problems that span service boundaries and coordinate with other teams when required.
Mentor engineers in distributed systems, networking, observability, reliability, and software engineering.
Conduct design and code reviews and raise engineering standards within the team.
Evaluate new technologies and engineering approaches where they improve scalability, reliability, performance, or operational efficiency.
Contribute to hiring, technical interviews, knowledge sharing, and development of engineering talent.
Skills & Technologies
The candidate should have strong expertise in several of the following areas:

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App