Live opening · Posted 17 days ago

Senior Network Engineer — AI Datacenter

Nava · Bengaluru, Karnataka, India (On-site)
Linkedin No
You are 17 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 17 days ago
CompanyNava
LocationBengaluru, Karnataka, India (On-site)
Work modeNo
SourceLinkedin
Listed17 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
4 min from Linkedin publishing this role to us finding it
14 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
63,206 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

About Nava
Nava is a neocloud company purpose-built for the AI era. We design, deploy, and operate large-scale GPU infrastructure and deliver inference-as-a-service to teams building the next generation of AI products. Our platform runs on NVIDIA GPU systems, high-performance RDMA fabrics, and a fully automated, software-defined operations model. Every engineer at Nava works close to the metal, on infrastructure built to keep GPUs saturated and models serving.
About The Team
The Network Engineering team owns the full lifecycle of Nava's AI datacenter fabric, from architecture and design through deployment, automation, and production operation. We build the lossless, high-throughput networks that carry GPU-to-GPU traffic for distributed training and low-latency inference. This is a hands-on team where the network is treated as code.
Responsibilities
Design and deliver leaf-spine fabrics using BGP and EVPN-VXLAN, and implement RoCEv2 and InfiniBand lossless networking for GPU backend traffic.
Build and extend network automation in Python, Ansible, and Terraform to provision, validate, and operate the fabric.
Lead deployment and turn-up of new GPU clusters, including acceptance testing and performance validation against line-rate targets.
Triage and resolve production network incidents; tune congestion control (PFC/ECN) to keep RDMA traffic lossless under real load.
Partner with Compute and Storage engineers to integrate the network end-to-end and eliminate bottlenecks.
Contribute to standards and mentor Network Engineers.
Required Qualifications
5–8 years in datacenter or production network engineering with strong hands-on BGP and EVPN-VXLAN experience.
Strong hands-on command of core networking protocols: BGP, OSPF, IS-IS, TCP/IP, IPv4 and IPv6, DNS, DHCP, and MPLS.
Experience with networking protocols such as TCP/IP, VPN, DNS, DHCP, and SSL/TLS.
Solid understanding of datacenter fabric concepts (leaf-spine, EVPN-VXLAN) and RDMA fabrics (RoCEv2 and InfiniBand), including lossless Ethernet.
Automation skills: Python plus Ansible and/or Terraform.
Comfortable operating in a fast-paced, on-call production environment.
Preferred Qualifications
Experience with NVIDIA networking (Spectrum-X, Quantum InfiniBand, BlueField DPUs) and NCCL traffic patterns.
Experience operating GPU clusters for large-scale distributed training or inference.
Familiarity with network telemetry, streaming analytics, and closed-loop automation.
Relevant certifications (e.g., CCNP or vendor equivalents).
Experience with Cisco and/or Arista platforms is a strong plus.
Technology environment
At Nava, Our Network Infrastructure Is Deeply Integrated With Modern AI Workloads—designed For Scale, Performance, And Automation. Below Is An Overview Of The Key Technologies And Platforms You’ll Work With Daily
Hardware & Interconnects: NVIDIA DGX/HGX systems (including Blackwell architecture), NVLink, Spectrum and Quantum-based switches, BlueField DPUs, and InfiniBand fabrics.
RDMA & Lossless Networking: RoCEv2 and InfiniBand RDMA, with deep familiarity in congestion control mechanisms (PFC, ECN) and lossless Ethernet configuration.
Control & Data Plane Protocols: BGP (including MP-BGP for EVPN), OSPF, IS-IS, EVPN-VXLAN for fabric virtualization, MPLS, and full IPv4/IPv6 stack support.
Automation & Tooling: Python for custom tooling, Ansible for configuration management, Terraform for infrastructure-as-code, and Linux-based toolchains (e.g., iproute2, netlink).
Observability & Operations: Telemetry via streaming telemetry (gNMI/sFlow), integration with monitoring stacks (Prometheus/Grafana), and experience with closed-loop automation for anomaly detection and remediation.
You’ll operate in a unified stack where networking, compute, and storage are co-designed—making deep technical understanding and automation-first thinking essential to success.
Skills: bgp,design,nvidia,rdma,networking,infiniband,python,ansible,ip,automation

Work arrangement
No

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App