Live opening · Posted 27 days ago

Software Engineer - Kubernetes / GPU

eBay · Bangalore
Instahyre 6-10 yrs
You are 27 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 27 days ago
CompanyeBay
LocationBangalore
Experience6-10 yrs
SourceInstahyre
Listed27 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
0 min from Instahyre publishing this role to us finding it
6 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
16,251 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

We are looking for an experienced Software Engineer specializing in Kubernetes and GPU infrastructure to design and operate the foundational systems that power eBay's AI platform. You will own critical layers of our infrastructure, from Kubernetes CRD-based automation and a custom AI-aware GPU scheduler to RDMA-optimized multi-NIC GPU clusters and large-scale training environments, enabling our ML teams to train and serve AI models at eBay scale. You will work on Kubernetes operator development, Gateway API and hybrid cloud networking, multi-NIC RDMA fabric design, a global GPU scheduler with cross-availability-zone dispatch, GPU pool management with provisioned throughput integration, topology-aware workload placement, and KubeRay infrastructure, partnering closely with ML Platform, AI Research, and Networking teams.
Responsibilities:
Design and build Kubernetes Custom Resource Definitions (CRDs) and operators for ML workloads, GPU node pools, and RayService CRDs.
Architect and operate Kubernetes networking layers, including Gateway API, Service Load Balancers, and API gateways for hybrid cloud connectivity.
Design, deploy, and operate multi-NIC Kubernetes clusters with RDMA-enabled networking.
Implement and optimize RDMA networking using GPUDirect RDMA, RoCE, and InfiniBand for distributed GPU workloads.
Build and operate a custom AI-aware GPU scheduler with topology-aware placement, preemption, and GPU defragmentation.
Design and manage GPU pool management systems across on-premise and cloud-burst GPU environments.
Deploy and operate KubeRay infrastructure for distributed Ray clusters supporting training and inference workloads.
Implement cloud bursting and spot or preemptible GPU scheduling to improve utilization.
Automate GPU asset provisioning, node configuration, and cluster lifecycle management using Infrastructure-as-Code and GitOps.
Implement networking policies, multi-tenant isolation, RBAC, and security controls across the Kubernetes environment.
Build observability for GPU utilization, NCCL communication, scheduler decisions, and network throughput.
Collaborate with the ML Platform, AI Research, and Networking teams to optimize infrastructure for training and online inference.

Experience
6-10 yrs

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App