Live opening · Posted 3 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Senior Kubernetes Developer
Location: Dallas, TX Preferred
Work Arrangement: Hybrid, 3 Days Onsite / 2 Days Remote
Remote Flexibility: Full remote may be considered for the right candidate
Relocation: Available for qualified non-local candidates
Employment Type: Direct Hire
Overview
Our client is seeking a Senior Kubernetes Developer to design and build the software powering a next-generation GPU-accelerated compute platform supporting AI, machine learning, LLM, and HPC workloads.
This is a software development role focused on Kubernetes, not a traditional DevOps, SRE, or Kubernetes administration position.
The core focus is building Kubernetes-native software including custom operators, controllers, CRDs, APIs, schedulers, and internal platform services used to orchestrate large-scale GPU infrastructure.
The ideal candidate is a strong developer who understands Kubernetes internals and has experience building software on top of Kubernetes, not simply deploying applications or maintaining clusters.
Key Responsibilities
Develop Kubernetes-native software using Go, Python, or similar languages.
Build custom operators, controllers, CRDs, APIs, and platform services.
Extend Kubernetes to support GPU-intensive AI/ML and HPC workloads.
Develop automation for cluster provisioning, lifecycle management, scheduling, and infrastructure orchestration.
Build GPU scheduling, allocation, workload placement, and resource-isolation capabilities.
Integrate NVIDIA technologies including GPU Operator, device plugins, MIG, and DCGM.
Develop internal tools and APIs for provisioning and managing GPU compute resources.
Improve platform scalability, GPU utilization, workload performance, and reliability.
Integrate Kubernetes with high-performance networking, storage, and bare-metal infrastructure.
Build observability and automated remediation capabilities for distributed compute environments.
Required Qualifications
Strong software development experience with Go, Python, or another modern programming language.
Hands-on experience building Kubernetes operators, controllers, CRDs, APIs, or similar Kubernetes-native software.
Strong understanding of Kubernetes architecture, controllers, reconciliation, scheduling, RBAC, networking, and cluster lifecycle.
Experience building platforms or distributed systems on Kubernetes.
Experience with GPU infrastructure and NVIDIA technologies.
Experience supporting AI/ML, LLM, HPC, or other compute-intensive workloads.
Strong Linux and distributed systems knowledge.
Experience with Terraform, Helm, Kustomize, Argo CD, Flux, or similar tooling.
Ability to troubleshoot across Kubernetes, compute, networking, storage, GPU, and application layers.
Preferred Experience
NVIDIA GPU clusters.
Slurm, Volcano, kube-scheduler extensions, or custom scheduling.
CUDA, NCCL, PyTorch, or TensorFlow.
InfiniBand, RDMA, RoCE, or other high-performance networking.
Bare-metal Kubernetes.
Internal developer platforms or self-service infrastructure.
AI infrastructure, HPC, cloud infrastructure, or large-scale distributed systems.
Ideal Candidate
The ideal candidate is a software developer who builds Kubernetes-native systems.
This person should be comfortable writing operators, controllers, APIs, schedulers, and automation that extend Kubernetes and manage complex GPU infrastructure.
Candidates whose background is primarily DevOps, CI/CD, Terraform administration, application deployment, or Kubernetes operations without substantial software development experience are unlikely to be the right fit.
Dallas-based candidates are preferred, but full remote may be considered for candidates with exceptional Kubernetes development and GPU infrastructure experience.
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.