Live opening · Posted 5 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Senior AI & HPC Infrastructure Engineer Remote (US) | $150,000 - $170,000 Base + Bonus
Our client is a highly respected global consulting and research organization that supports leading commercial enterprises, public sector institutions, and professional services firms. As part of a significant investment in AI infrastructure, they are expanding their internal high performance computing (HPC) and GPU capabilities to support next-generation machine learning and large language model (LLM) initiatives.
This is an exciting opportunity to join a small, highly skilled infrastructure team at a pivotal stage of growth. The successful candidate will play a key role in scaling a GPU environment from 8 to 32 NVIDIA H200 GPUs while helping shape the organization's long-term AI and HPC strategy.
The role is heavily project-focused, with approximately 85% dedicated to engineering, architecture, and platform development activities, and a smaller proportion supporting operational needs.
The Opportunity
You'll work at the intersection of traditional HPC, research computing, and modern AI infrastructure, supporting both analytical workloads and large-scale model training environments.
Key responsibilities include:
Designing, deploying, and maintaining GPU-accelerated computing infrastructure
Supporting large-scale AI/ML and LLM training environments
Managing Linux-based HPC clusters and associated services
Administering and optimizing parallel file systems, particularly IBM Spectrum Scale (GPFS)
Managing NVIDIA GPU platforms, CUDA, cuDNN, NCCL, and related tooling
Supporting resource scheduling through SLURM and related technologies
Performance tuning for distributed and multi-GPU workloads
Building automation, monitoring, reporting, and operational tooling
Collaborating with researchers, data scientists, and technical stakeholders to translate business requirements into infrastructure solutions
Evaluating emerging AI infrastructure technologies and recommending future platform enhancements
Required Experience
We're particularly interested in candidates who combine deep infrastructure expertise with an understanding of how end-users consume HPC resources.
Essential Skills
Strong Linux systems administration background
Experience supporting HPC or research computing environments
Hands-on NVIDIA GPU infrastructure experience
Strong experience with GPFS / IBM Spectrum Scale
Familiarity with job schedulers such as SLURM, LSF, or similar
Experience supporting distributed compute environments
Ability to lead projects independently with minimal oversight
Strong troubleshooting experience across compute, storage, networking, and hardware layers
Excellent communication skills with the ability to explain complex technical concepts to non-technical stakeholders
Highly Desirable
Experience supporting AI/ML infrastructure or LLM platforms
Kubernetes and container orchestration experience
MLOps tooling exposure (MLflow, Kubeflow, etc.)
Experience tuning large model training and inference workloads
Bright Cluster Manager
Ansible
Docker, Apptainer, or Singularity
SAS platform experience alongside GPU/HPC expertise
Ideal Background
This role is particularly well suited to someone who:
Has approximately 5-10 years of relevant infrastructure experience
Is currently operating at a strong mid-level and ready for a senior step forward
Can own and deliver significant technical projects independently
Comes from an HPC, research computing, higher education, government, scientific computing, or technical enterprise environment
Enjoys balancing infrastructure engineering with emerging AI technologies
Team & Culture
You'll join a close-knit team of experienced infrastructure specialists covering HPC operations, automation, applications, and platform engineering. The group operates in a highly collaborative remote-first model, with team members distributed across the United States.
A notable area of planned growth is AI and LLM infrastructure expertise, making this a high-visibility hire with significant opportunity for progression and influence.
If you're an HPC infrastructure engineer looking to move into a highly visible AI-focused environment while remaining close to the hardware, architecture, and platform engineering side of the stack, we'd love to hear from you.
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.