Failure Analysis Engineer
Hyve Solutions Corporation
Apply to this jobHyve Company Overview
Hyve Solutions transforms complex engineering challenges into production reality for technology innovators building AI, cloud, and connected infrastructure. As a US-based manufacturing partner, the company rapidly delivers fast, agile execution through dep technical partnerships and co-innovation, and supply chain clarity. Hyve’s integrated ODM, CM, and SI capabilities eliminate vendor complexity while accelerating time-to-market with single-partner accountability from design through scale. The company co-innovates with deep engineering expertise, treating customer success as its own while building tomorrow's digital infrastructure.
Position Summary
We are seeking a highly analytical Failure Analysis Engineer to support the investigation of hardware failures in rack systems, server platforms, and data center infrastructure products. This role is responsible for diagnosing complex electrical, mechanical, thermal, and system-level failures throughout the product lifecycle, including manufacturing, qualification, and customer returns.
Key Responsibilities
Perform failure isolation at the subsystem and rack level, as well as root cause analysis on failures involving server systems, rack-level assemblies, storage platforms, networking hardware, and associated components.
Investigate failures from manufacturing, system integration, and customer returns (RMA).
Analyze electrical, mechanical, thermal, and firmware-related failures using structured troubleshooting methodologies.
Utilize laboratory equipment including oscilloscopes, digital multimeters, and optical microscopes.
Conduct system-level debugging of server motherboards, backplanes, power distribution boards (PDBs), power supplies, GPU modules, CPUs, DIMMs, NICs, storage devices, and PCIe components.
Work closely with Design Engineering, Manufacturing Engineering, Quality, Reliability, Supplier Quality, Test Engineering, and Operations to identify corrective actions.
Lead Root Cause Analysis (RCA) activities using 8D, 5-Why, Fishbone Diagram, Fault Tree Analysis (FTA), and Failure Modes and Effects Analysis (FMEA).
Develop and publish detailed failure analysis reports, including technical findings, corrective actions, and preventive recommendations.
Identify recurring failure trends through statistical analysis and recommend design or process improvements.
Drive corrective and preventive actions (CAPA) to improve product reliability and manufacturing yield.
Collaborate with suppliers to investigate component-level failures and improve incoming material quality.
Support customer escalations by providing technical expertise during failure investigations.
Required Qualifications
Bachelor’s degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, or a related engineering discipline.
3+ years of experience in failure analysis, hardware validation, quality engineering, reliability engineering, or manufacturing engineering.
Experience supporting enterprise servers, rack systems, storage platforms, networking equipment, or data center infrastructure.
Strong understanding of server architecture including, but not limited to CPUs, GPUs, Memory (DDR4/DDR5), PCIe architecture, NVMe storage, Ethernet networking, BMC/IPMI management and Power distribution systems,
Experience troubleshooting complex hardware failures at the system level.
Knowledge of schematic review and hardware debugging techniques.
Ability to interpret manufacturing and test logs to identify failure mechanisms.
Excellent analytical, communication, and technical documentation skills.
Key Competencies
Strong analytical and troubleshooting skills
Cross-functional collaboration
Data-driven decision making
Technical writing and presentation
Continuous improvement mindset
Ability to manage multiple high-priority investigations in a fast-paced environment
The anticipated base salary range for this position is $142,000-$169,000 annually.
Actual compensation will be determined based on experience, skills, qualifications, and business needs.
What’s in it for you
Benefit Insurance
- Flexible Spending Account (FSA)
- Health Savings Account (HSA)
- Mental Health Care
- 401K with match
- Paid Holidays, Vacation & Sick Days
- Tuition Reimbursement
- LEAP Program
- MyFlexPay
We are an Equal Opportunity Employer and do not discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, protected veteran status, disability, or any other legally protected status. We are committed to creating an inclusive environment for all employees.
Summary
Investigate hardware failures in server and data center products, develop corrective actions.
Job title
Failure Analysis Engineer
Experience level
3+ years
Minimum experience
3+ years exp
Industry
manufacturing
Location requirements
Fremont, CA; remote work not specified.
Salary
$142k–$169k
Management role
No
Required skills
Preferred skills
Specializations
Structured locations inferred from the posting.
Fremont, CA, USA