HPC Engineer
Institute of Foundation Models
Apply to this job Sunnyvale, CA on site Until 9/22/2026 2+ years exp First posted June 1, 2026 Last posted July 24, 2026
Job description
About MBZUAI
The Institute for Foundation Models (IFM) operates some of the world's largest AI supercomputing environments.
Position Summary
This role provides operational coverage during Abu Dhabi overnight hours and serves as a primary point of contact for infrastructure monitoring, incident triage, researcher support, and production operations.
The Institute for Foundation Models (IFM) operates some of the world's largest AI supercomputing environments.
Position Summary
This role provides operational coverage during Abu Dhabi overnight hours and serves as a primary point of contact for infrastructure monitoring, incident triage, researcher support, and production operations.
Responsibilities
• Monitor health, performance, and availability of large-scale GPU clusters.
• Respond to incidents and perform first-level triage.
• Support researchers and troubleshoot job failures.
• Execute operational runbooks and recovery procedures.
• Validate cluster deployments, upgrades, and maintenance activities.
• Track infrastructure utilization and operational metrics.
• Develop automation and monitoring tools.
• Contribute to documentation and reporting.
• Respond to incidents and perform first-level triage.
• Support researchers and troubleshoot job failures.
• Execute operational runbooks and recovery procedures.
• Validate cluster deployments, upgrades, and maintenance activities.
• Track infrastructure utilization and operational metrics.
• Develop automation and monitoring tools.
• Contribute to documentation and reporting.
Education
Bachelor's degree in Computer Science, Computer Engineering, Software Engineering, Information Technology, Electrical Engineering, Mathematics, Physics, or related disciplines.
Experience
• 2+ years in Linux systems administration, SRE, DevOps, cloud operations, HPC, or infrastructure operations.
• Strong Linux troubleshooting skills.
• Experience with scripting using Python or Bash.
• Strong Linux troubleshooting skills.
• Experience with scripting using Python or Bash.
Preferred Qualifications
• Slurm.
• GPU infrastructure.
• AWS, Azure, or GCP.
• Grafana, Prometheus, Datadog, or similar tools.
• Containers and Kubernetes.
• AI/ML infrastructure exposure.
• Research computing environments.
• GPU infrastructure.
• AWS, Azure, or GCP.
• Grafana, Prometheus, Datadog, or similar tools.
• Containers and Kubernetes.
• AI/ML infrastructure exposure.
• Research computing environments.
Benefits Include
*Comprehensive medical, dental, and vision benefits
*Bonus
*401K Plan
*Generous paid time off, sick leave and holidays
*Paid Parental Leave
*Employee Assistance Program
*Life insurance and disability
Compensation (from employer):
150000–300000 USD per year
About this role
Summary
Monitor GPU clusters, support research, troubleshoot, automate, and optimize HPC infrastructure
Job title
HPC Engineer
Experience level
2+ years
Minimum experience
2+ years exp
Industry
software
Location requirements
onsite in Sunnyvale, CA with remote possible
Salary
$150k–$300k
Management role
No
Skills & keywords
Required skills
Linux troubleshootingPythonBash
Preferred skills
SlurmGPU infrastructureAWSAzureGCPGrafanaPrometheusDatadogcontainersKubernetesAI/ML infrastructure
Specializations
GPU clustersLinuxHPCDevOpscloud
Locations
Structured locations inferred from the posting.
Sunnyvale, CA, USA
Hybrid City
Related searches