Live opening · Posted 2 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Senior DevOps Engineer – HPC / EDA / SLURM / Azure
Location: Remote / Hybrid – Rancho Cordova, CA
Duration: 12 Months Contract
Department: IT Datacenter Infrastructure (ITDC)
Job Overview
We are looking for an experienced Senior DevOps Engineer to support a High Performance Computing (HPC) and Electronic Design Automation (EDA) infrastructure environment.
The ideal candidate will have strong hands-on experience with Linux, HPC/SLURM, Ansible, Terraform, Azure, and enterprise authentication. This person should be comfortable working independently, troubleshooting production systems, automating infrastructure, and coordinating with multiple technical teams.
Key Responsibilities
Administer and support Linux-based HPC environments.
Manage and support SLURM clusters, including compute nodes, partitions, configurations, and migrations.
Develop and maintain Ansible playbooks and roles for Linux server configuration and automation.
Use Terraform for infrastructure provisioning and configuration changes.
Support Azure cloud infrastructure used for EDA/HPC workloads.
Support SLES 15 and Ubuntu Linux environments.
Configure and troubleshoot Linux authentication using SSSD, LDAP, Active Directory, and Okta.
Manage Linux user/group access, including UID/GID troubleshooting and reconciliation.
Support enterprise storage environments such as NetApp, NFS, and AutoFS.
Troubleshoot Linux services and infrastructure issues involving VNC/ThinLinc, AutoFS, Datadog, and other platform services.
Support logging and monitoring solutions such as Splunk and Datadog.
Use Git/GitHub for infrastructure code, pull requests, code reviews, and repository management.
Work with Artifactory for configuration artifacts and software packages.
Follow ServiceNow change management processes for production changes.
Create and maintain MOPs, runbooks, technical documentation, and architecture diagrams.
Collaborate with EDA, storage, security/identity, and infrastructure teams.
Required Skills
Candidates should have strong experience in most of the following:
9+ years of experience in DevOps, Platform Engineering, Linux Systems Engineering, or similar roles.
Hands-on HPC administration experience.
Strong experience with SLURM or another HPC workload manager.
Strong Linux administration skills.
Experience with SLES 15 and/or Ubuntu.
Strong Ansible experience, including playbooks and roles.
Experience with Terraform / Infrastructure as Code (IaC).
Experience with Azure cloud environments.
Experience with SSSD, LDAP, Active Directory, or Okta.
Experience with NFS, NetApp, AutoFS, or similar enterprise storage technologies.
Strong Bash and/or Python scripting skills.
Experience with Git/GitHub.
Experience with monitoring/logging tools such as Splunk or Datadog.
Ability to troubleshoot production Linux and infrastructure issues.
Strong technical documentation and communication skills.
Preferred Experience
Experience supporting EDA, semiconductor, scientific computing, or HPC environments.
Experience with SUSE Linux Enterprise Server (SLES).
Experience with ThinLinc / VNC.
Experience with NetApp storage in HPC environments.
Experience with Artifactory.
Experience with ServiceNow change management.
Experience with SLURM cluster migrations or datacenter migrations.
Experience working in semiconductor, storage, or high-tech companies.
Ideal Candidate Profile
We are especially interested in candidates who are:
Linux Engineer + HPC Engineer + DevOps Engineer
with hands-on experience in:
SLURM + Ansible + Terraform + Azure + SLES/Ubuntu + SSSD/LDAP/AD + NetApp/NFS
Candidates should be comfortable working in a production environment and handling infrastructure automation, troubleshooting, migrations, and cross-team coordination.
Contract Details
Position: Senior DevOps Engineer – HPC / EDA
Contract: 12 Months
Location: Remote / Hybrid – Rancho Cordova, CA
Experience: 5+ Years
Work Environment: Enterprise HPC / EDA / IT Datacenter Infrastructure
Interested candidates: Please share your updated resume highlighting your experience with HPC, SLURM, Linux, Ansible, Terraform, and Azure.
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.