Live opening · Posted 4 days ago

Senior Engineer, Cloud Infrastructure and Networking

Skylo · United States (Remote)
Linkedin Yes
You are 4 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 4 days ago
CompanySkylo
LocationUnited States (Remote)
SalaryMedical benefit
Work modeYes
SourceLinkedin
Listed4 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
11 min from Linkedin publishing this role to us finding it
14 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
63,223 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

The world still has coverage blind spots. You could help eliminate them at Skylo.
Skylo has pioneered a standards-based approach to satellite connectivity. We connect smartphones and IoT devices directly to satellites. No special hardware, no entirely new networks. Just billions of existing devices, suddenly reachable anywhere on Earth. We're not building toward this future. We're already in it.
Our direct-to-device service is live on millions of activated devices across five continents, covering more than 72 million square kilometers, in partnership with leading satellite operators, mobile network operators, Tier-1 chipset makers, and OEMs worldwide. And we're just getting started.
At the heart of it all is Skylo's commercial NTN vRAN: a 3GPP standards-based, cloud-native platform that seamlessly bridges terrestrial and satellite networks. It's the infrastructure that makes true anywhere, anytime connectivity possible.
When you join Skylo, you'll work at the intersection of three markets reshaping how the world stays connected: mass-market consumer devices, automotive, and industrial IoT. Enabling people outdoors and critical workflows in the world's most remote places.
This is a rare chance to work on technology that matters, at a company that's already proving it works
About Skylo
Skylo is a global Non-Terrestrial Network (NTN) service provider based in Mountain View, CA, offering a service that allows smartphone and IoT cellular devices to connect directly over existing satellites.
Skylo's direct-to-device service is live on millions of activated devices across five continents, with more than 60 million square kilometers of coverage, in partnership with multiple satellite operators, mobile network operators (MNOs), Tier-1 chipset makers, and OEMs. Devices connected over satellite are managed and served by Skylo's commercial NTN vRAN — a 3GPP standards-based, cloud-native base station and core. Skylo provides an anywhere, anytime connectivity solution that seamlessly roams between terrestrial and satellite networks. Our focus is on enabling connected services across three main verticals: mass-market consumer devices, automotive, and industrial IoT.
How You Will Impact Skylo
As a Senior, Cloud Infrastructure and Networking, in the Global Product Support & Customer Success organization, you are the Cloud Infrastructure domain authority within Skylo's production NTN network. Everything runs on the infrastructure you keep healthy — RAN NFs, Core NFs, OSS, BSS, and the observability pipeline itself. When a GKE node fails, when ArgoCD drifts, when a Persistent Volume Claim goes unavailable, when a PostgreSQL replica falls behind, when Prometheus WAL corrupts — you own the response.
You operate across Skylo's full hybrid cloud estate: GCP public cloud (GKE clusters, Pub/Sub pipelines, Cloud SQL) and on-premise private cloud infrastructure (bare-metal Kubernetes, hyperconverged compute, software-defined storage). You own 24x7 platform health, the observability pipeline (Prometheus, VictoriaMetrics, Grafana, OpenTelemetry), persistent storage operations (PostgreSQL, Redis), and the operational interface with Network Implementation for all GitOps-driven infrastructure changes.
Key Responsibilities
Cloud Infrastructure Operations & Health Ownership
Own 24x7 cloud infrastructure health across Skylo's hybrid production environment: GKE cluster node status, namespace and pod health, Persistent Volume Claim availability, network policies, and multi-cluster federation across Skylo's GCP footprint.
Own on-premise Kubernetes cluster health: bare-metal node availability, container runtime stability, CNI networking, persistent storage arrays (Ceph/Rook or equivalent), and hyperconverged compute platform operations (Harvester, KubeVirt, or KVM).
Monitor and triage infrastructure alarms using OSS dashboards, Grafana/VictoriaMetrics telemetry, GCP Cloud Monitoring, and Loki log correlation — distinguish transient platform events from systemic infrastructure degradation.
Execute and own Cloud Infra runbooks for P2–P4 fault categories: GKE node recovery, pod eviction and rescheduling, PVC repair, database failover execution, Prometheus WAL corruption recovery, ArgoCD drift remediation, and certificate rotation — without requiring engineering involvement for covered fault classes.
Own BSS-IIS GKE cluster monitoring and infrastructure health; maintain runbooks that reflect current cluster topology after every infrastructure change.
Observability Pipeline & Data Platform Operations
Own the observability pipeline end-to-end: Prometheus scrape target integrity, VictoriaMetrics retention and query performance, Grafana dashboard coverage and accuracy, OpenTelemetry collector health, and alert routing via Pub/Sub to the OSS.
Maintain database reliability: PostgreSQL streaming replication health, backup and restore procedures, failover testing, query performance monitoring; Redis cluster operations, eviction policy management, and persistence configuration.
Ensure log aggregation pipeline health (Loki or ELK): ingestion rates, retention policies, query performance, and completeness — the observability stack must be operational before the network events it monitors can be triaged.
Partner with NI (Network Implementation & Infrastructure) on all planned infrastructure changes: receive advance notice, validate post-deployment observability, and sign off on operational readiness before the change window closes.
Cloud Infrastructure Incident Diagnosis & Escalation Authority
Serve as the L3 escalation authority for all Cloud Infra incidents: take ownership from the Incident Manager, diagnose at the Kubernetes, storage, network, and database layer using kubectl, GCP console, node logs, and infrastructure telemetry, and deliver a resolution or a decision-grade root cause.
Lead Cloud Infra troubleshooting bridges: command the technical investigation for GKE node failures, cluster upgrade failures, storage outages, PubSub pipeline disruptions, database failover events, and ArgoCD sync failures — drive to resolution or clear engineering handoff.
Diagnose and resolve infrastructure failure modes: node NotReady conditions, pod CrashLoopBackOff chains, PVC mount failures, CSI driver errors, network policy misconfigurations, Helm release drift, etcd latency spikes, and cross-cluster federation breaks.
Participate in the global 24x7 on-call rotation as the Cloud Infra domain escalation tier — reachable within defined SLA windows for Sev 1 events; function as the technical decision-maker, not the first responder.
SLO Engineering & Reliability
Define and maintain SLOs for all Cloud Infra components: GKE control plane availability, database query latency, storage IOPS, message pipeline throughput, and observability stack uptime — tied directly to network SLA commitments to MNO partners.
Own error budget tracking and the process for trading error budget against deployment velocity; escalate when error budget burn rate requires engineering intervention or deployment freezes.
Drive toil reduction: identify and eliminate manual Cloud Infra procedures; own the roadmap to automated cluster recovery, rolling restarts, storage repair, and certificate rotation in partnership with Ops Platform Engineering.
Lead capacity planning for compute, storage, and network resources across public and private cloud — forecast growth based on subscriber projections and new MNO partner onboarding.
Root Cause Analysis & Post-Incident Ownership
Own Cloud Infra RCA end-to-end: lead the investigation, document the complete causal chain from infrastructure trigger through upstream NF impact, and deliver systemic action items with owners, timelines, and measurable success criteria.
Deliver Initial RCA documentation within defined SLA windows; identify systemic infrastructure failure patterns — cluster upgrade regressions, storage controller bugs, network policy drift, resource exhaustion trends — and translate them into engineering requirements.
Contribute to the weekly and monthly Network Performance Report: infrastructure availability, database latency trends, storage IOPS, observability pipeline health, and SLA deviation analysis.
Runbook Authorship & Operational Standards
Author, own, and maintain all Cloud Infra runbooks and SOPs: GKE node recovery, database failover, storage expansion, Prometheus WAL repair, ArgoCD rollback, certificate rotation, and cluster upgrade procedures — every procedure tested before production reliance.
Define the diagnostic decision tree for each known infrastructure fault class: entry condition, triage steps, isolation method, resolution action, and escalation criteria — written at the level where a Senior NRE can execute independently.
Validate and sign off on operational readiness for all infrastructure changes: DCI/DCE build-outs, GKE cluster expansions, Kubernetes version upgrades, and new on-premise hardware deployments.
Cross-Functional Collaboration & Team Development
Partner with OPE as the Cloud Infra domain's primary automation consumer: define Kubernetes event schemas, alert-to-action contracts, and closed-loop policy requirements for infrastructure auto-remediation.
Represent Cloud Infra in NI architecture reviews: define observability and operational readiness requirements for all infrastructure expansions and GitOps pipeline changes before go-live.
Collaborate with Core NRE and R

Work arrangement
Yes

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App