Live opening · Posted 27 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Requirements:
5+ years of experience building distributed systems or infrastructure platforms with deep Kubernetes expertise.
Strong programming skills in Go and/or Python.
Familiarity with Kubernetes controller development frameworks such as Kubebuilder, Operator SDK, or controller-runtime.
Deep understanding of Kubernetes internals, including the API server, etcd, scheduler, controller manager, and kubelet.
Hands-on experience designing Kubernetes networking, including Gateway API, CNI plugins, service load balancing, and hybrid cloud architectures.
Experience designing and operating multi-NIC Kubernetes clusters using NVIDIA Network Operator, SR-IOV device plugins, or equivalent tooling.
Strong understanding of RDMA networking protocols and NCCL configuration for distributed GPU workloads.
Experience with Kubernetes GPU scheduler frameworks and GPU pool management, including MIG partitioning and preemption policies.
Hands-on experience deploying and operating KubeRay for distributed Ray workloads.
Experience with GPU asset lifecycle management, bare-metal provisioning automation, and GitOps-based CD tooling such as Argo CD.
Familiarity with GPU technologies including NVIDIA CUDA, NVLink, NVSwitch, device plugins, and the GPU Operator ecosystem.
Experience with observability tooling such as Prometheus, Grafana, and OpenTelemetry.
Strong debugging and performance optimization skills across GPU driver stacks, RDMA networking, and distributed Kubernetes infrastructure.
Experience
5-9 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.