Live opening · Posted 7 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
This is a fully remote role open to candidates across the MENA region
About Saturn Cloud
Saturn Cloud builds infrastructure for running AI, machine learning, and data workloads at scale. Our platform helps teams develop, deploy, and operate compute-intensive workloads across modern cloud and GPU infrastructure.
We’re looking for a GPU/HPC-focused Site Reliability Engineer based in the MENA region to help operate and troubleshoot the large-scale GPU infrastructure supporting Saturn Cloud Token Factory.
This role is focused on the infrastructure below and around the Kubernetes layer. You’ll serve as a technical escalation point for GPU health, NVIDIA software, high-performance networking, topology, and distributed GPU performance issues affecting production inference workloads.
The ideal candidate comes from GPU infrastructure, HPC, AI infrastructure, neocloud, hyperscaler, or large-scale ML platform operations rather than traditional application support.
What You’ll Do
Diagnose and resolve production issues affecting large-scale GPU inference infrastructure
Troubleshoot NVIDIA datacenter GPUs, drivers, CUDA compatibility, and GPU container runtimes
Investigate GPU health issues, including Xid errors and hardware or driver failure modes
Diagnose PCIe, NUMA, GPU placement, NVLink, and NVSwitch issues
Troubleshoot multi-GPU and multi-node workloads
Investigate high-performance networking and distributed communication failures
Distinguish application and inference-runtime issues from GPU, fabric, topology, node, driver, or hardware failures
Work with Kubernetes-based GPU workloads and NVIDIA GPU Operator/device plugins
Use production observability and GPU metrics to diagnose reliability and performance issues
Work directly with GPU-cloud and infrastructure providers when incidents require hardware or fabric investigation
Produce clear technical evidence showing where the failure occurs and what the appropriate infrastructure team needs to investigate
What We’re Looking ForCore Skills
Deep Linux systems debugging experience
NVIDIA datacenter GPU administration
NVIDIA driver installation, upgrades, and troubleshooting
Strong understanding of CUDA and driver compatibility
NVIDIA Container Toolkit/runtime
Experience with NVML and nvidia-smi
GPU health diagnostics, including Xid errors and common hardware/driver failure modes
Understanding of PCIe topology, NUMA, and GPU placement
Understanding of NVLink and NVSwitch fundamentals
Familiarity with Kubernetes GPU Operator and device plugins
Experience running containerized GPU workloads
High-Performance Networking
You should have strong experience with several of the following:
InfiniBand
RDMA
RoCE
NCCL
GPUDirect RDMA
NIC/GPU topology
NCCL testing and distributed workload diagnostics
Bandwidth and latency troubleshooting
Multi-node GPU communication failures
You should be able to determine whether a production issue originates in the application, GPU, network fabric, topology, node, or another underlying infrastructure layer.
Inference Infrastructure
You don’t need to be an ML researcher, but you should understand how modern inference workloads exercise GPU infrastructure.
Relevant experience includes:
vLLM, NVIDIA Dynamo, Triton, or comparable inference runtimes
Tensor and pipeline parallelism
Model loading and GPU memory consumption
KV cache
Continuous batching
GPU and memory utilization
Out-of-memory diagnosis
Multi-GPU and multi-node inference
Basic inference performance analysis, including throughput, latency, and GPU saturation
Additional Skills
Kubernetes troubleshooting sufficient to independently investigate GPU workloads inside a cluster
Prometheus/Grafana and NVIDIA DCGM metrics
Bash and Python
containerd/Docker
Experience with bare-metal GPU clusters or GPU cloud infrastructure
Familiarity with B200/B300/H200/H100-class systems is highly desirable
Strong Pluses
NVIDIA DCGM
NVIDIA Dynamo
KAI Scheduler or Grove
Spectrum-X
Mellanox/ConnectX networking
OFED/DOCA
Slurm or HPC cluster administration
Kubernetes-based GPU clouds
Large GPU fleet operations
GPU burn-in, qualification, and health-check tooling
What Success Looks Like
When an inference workload becomes unavailable or significantly slower, you can determine whether the problem is caused by the inference runtime, GPU memory pressure, a failed GPU, a driver/kernel interaction, PCIe/NVLink/NVSwitch topology, NCCL, RDMA/InfiniBand fabric, or the underlying node.
You can collect enough evidence to distinguish a Saturn Cloud software issue from an operator infrastructure or hardware problem and communicate precisely what an infrastructure provider needs to investigate.
You’re the person the team turns to when Kubernetes looks healthy, but the GPU workload isn’t.
Why Saturn Cloud
Remote-first culture with a high-trust, high-ownership environment
Work directly with cutting-edge GPU and AI infrastructure
Solve complex production problems spanning hardware, networking, Kubernetes, and modern inference systems
Help shape the reliability and operational practices behind large-scale AI workloads
Compensation & Benefits
Competitive salary
Flexible PTO
Fully remote position
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.