Live opening · Posted 1 day ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Responsibilities:
Deploy and manage ML/LLM models (< 20B parameters) in production environments.
Design and execute model load testing and performance benchmarking (latency, throughput, memory, cost).
Build and optimise multi-node, multi-GPU training pipelines.
Configure and tune distributed training frameworks (data parallelism, model parallelism, pipeline parallelism).
Optimise GPU utilisation, memory footprint, and inference costs.
Set up CI/CD pipelines for model deployment and retraining.
Troubleshoot GPU, networking, and performance bottlenecks.
Work across cloud platforms to ensure portability and vendor-agnostic deployments.
Requirements:
Strong experience with GPU workloads (NVIDIA GPUs, CUDA concepts).
Proven expertise in model deployment on AWS, Azure, and GCP.
Hands-on experience deploying models up to 20B parameters.
Experience with distributed training (multi-node, multi-GPU setups).
Deep understanding of load testing, stress testing, and benchmarking ML systems.
Experience
4-8 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.