System Engineer

Jobs.supermicro.com

Apply to this job
Until 9/28/2026 3+ years exp First posted July 30, 2026 Last posted July 30, 2026
Job description

Job Req ID: 29579

About Supermicro:

Supermicro® is a Top Tier provider of advanced server, storage, and networking solutions for Data Center, Cloud Computing, Enterprise IT, Hadoop/ Big Data, Hyperscale, HPC and IoT/Embedded customers worldwide. We are the #5 fastest growing company among the Silicon Valley Top 50 technology firms. Our unprecedented global expansion has provided us with the opportunity to offer a large number of new positions to the technology community. We seek talented, passionate, and committed engineers, technologists, and business leaders to join us.
 

Essential Duties and Responsibilities:

Includes the following essential duties and responsibilities (other duties may also be assigned):
• Familiar with the day-to-day operational support for Cluster, Storage, HPC, AI, Data Center and Cloud infrastructures.
• Builds Cluster, Storage, HPC, AI, Data Center and Cloud infrastructures in-house and onsite testing, deployment, and platforms accordingly to meet customer's requirement.
• Troubleshoot hardware and software issues in rack cabinet. Provide fixes in a timely manner.
• Documents complex test procedures and troubleshooting procedures related to servers/networks/clusters software and hardware.
• Familiar with Intel/AMD/NVIDIA development toolkits like CUDA, oneAPI, ROCm.
• Conduct tests and benchmarks against server hardware, storage, network, applications, HPC and AI/ML/DL workflows.
• Conduct performance testing and benchmarking for servers, GPUs, and HPC environments.
• Analyze results to identify bottlenecks and optimize system performance for AI/ML workloads.
• Design and configure high-speed network topologies (InfiniBand, Ethernet) for AI clusters.
• Configure network components to ensure optimal performance.
• Write Python scripts to automate testing, monitoring, and system optimization.
• Understanding of AI/ML frameworks (e.g., PyTorch, TensorFlow) and deployment requirements for LLMs.
• Monitor network health and server performance, proactively identifying and resolving issues.
• Programming experience with web applications, including frontend or backend.
• Collect, visualize, and analyze test and benchmark results.
• Programming experience with Python, Ansible and Linux shell scripting.
• Maintain and develop in-house and on-site automated test programs.
• Write technical documentation including test reports and standard operating procedure (SOP).

Qualifications:

• Bachelor's  or Master degree in Computer Science or equivalent work experience preferred.
• 3+ years of proven experience in a HPC/AI or Cloud/Network management.
• In-depth knowledge of Cloud/HPC/AI deployment and testing.
• Strong problem-solving and decision-making abilities, with a proactive approach to identifying and resolving issues.
• Excellent communication skills, both verbal and written, with the ability to collaborate and build strong relationships with stakeholders at all levels.
• It's a plus if you have CCNA/CCNP certificates.
• Positive attitude, desire to learn, time management, and strong interpersonal skills are a plus!

EEO Statement

Supermicro is an Equal Opportunity Employer and embraces diversity in our employee population. It is the policy of Supermicro to provide equal opportunity to all qualified applicants and employees without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, protected veteran status or special disabled veteran, marital status, pregnancy, genetic information, or any other legally protected status.

About this role

Summary

Support and optimize HPC, AI, cloud infrastructures, troubleshoot hardware/software, automate tasks.

Job title

System Engineer

Experience level

3+ years

Minimum experience

3+ years exp

Industry

technology

Location requirements

Remote work possible, location unspecified

Salary

Not specified

Management role

No

Skills & keywords

Required skills

pythonlinux shell scriptingnetworkingtroubleshootingCUDA

Preferred skills

CCNACCNPmachine learningperformance benchmarking

Specializations

HPCAIcloudnetworkingpython
Locations

Structured locations inferred from the posting.

No structured locations extracted for this role yet.