Live opening · Posted 11 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Role Overview
We are seeking seasoned professionals with deep expertise in operating and managing High-Performance Computing (HPC) platforms. The ideal candidate will have hands-on experience in designing, deploying, and maintaining HPC clusters, storage systems, and networking infrastructure, leveraging industry-leading tools and technologies.
Key Responsibilities
HPC Infrastructure Management
Operate and maintain HPC clusters based on CentOS, RHEL, and hardware platforms like HPE and NVIDIA DGX.
Ensure optimal performance, scalability, and reliability of compute resources.
Storage Administration
Manage large-scale storage systems including Dell Isilon, VAST Storage, Lustre, and GPFS.
Implement data lifecycle management and optimize storage performance for HPC workloads.
Networking
Configure and maintain InfiniBand-based networking for low-latency, high-bandwidth communication.
Troubleshoot network performance issues and ensure secure connectivity.
Cluster and Job Scheduling
Administer cluster management tools such as Bright Cluster Manager, Altair Grid Manager, and IBM LSF.
Optimize job scheduling and resource allocation for diverse workloads.
Monitoring and Automation
Implement monitoring solutions using Zabbix, Grafana, and ELK Stack.
Automate provisioning and configuration using Cobbler, Chef, Ansible, and AWS ParallelCluster.
Performance Tuning & Troubleshooting
Conduct performance benchmarking and tuning for HPC workloads.
Diagnose and resolve hardware/software issues across compute, storage, and network layers.
Security & Compliance
Ensure HPC environment adheres to security best practices and compliance standards.
Required Skills & Qualifications
Technical Expertise
Strong knowledge of Linux OS (CentOS, RHEL) and HPC hardware platforms (HPE, NVIDIA DGX).
Hands-on experience with parallel file systems (Lustre, GPFS) and enterprise storage solutions.
Proficiency in InfiniBand networking and high-speed interconnects.
Familiarity with job schedulers and cluster management tools (IBM LSF, Bright Cluster Manager, Altair Grid Manager).
Automation & Scripting
Expertise in Ansible, Chef, Cobbler, and scripting languages (Bash, Python).
Experience with AWS ParallelCluster or similar cloud-based HPC solutions.
Monitoring & Logging
Practical experience with Zabbix, Grafana, and ELK Stack for system health and performance monitoring.
Soft Skills
Strong problem-solving and analytical skills.
Ability to work in a fast-paced environment and lead technical teams.
Excellent communication and documentation skills.
Preferred Qualifications
Exposure to AI/ML workloads on HPC clusters.
Experience with containerization (Docker, Singularity) in HPC environments.
Knowledge of security hardening for HPC systems.
Education
Bachelor’s or Master’s degree in Computer Science, Engineering, or related field.
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.