Reliability Engineer

Intel Corporation

Apply to this job
US, Massachusetts, Beaver Brook US, California, Santa Clara Until 10/3/2026 4+ years exp First posted August 4, 2026 Last posted August 4, 2026
Job description

Job Details:

Job Description: 

Join us to help build the next generation of AI hardware solutions. You will be part of a highly skilled, agile team developing cutting-edge hardware for the AI domain, where we push the boundaries of what silicon can do for emerging AI workloads. With a startup-like culture, we move quickly and give engineers the opportunity to drive significant technical and business impact. 

We are continuously developing modern and effective working methods, including hands-on adoption of AI tools throughout the chip development flow.  

Mission: Define and own the pod-level reliability specifications that ensure the availability, resilience, and serviceability of a large-scale data center across hardware, thermal, and operational dimensions.

Responsibilities:

  • Define and maintain pod-level reliability/availability specs and targets (MTBF, AFR, RAS) for compute, memory, storage, network, power, and cooling subsystems.

  • Translate system/SLA requirements into pod and subsystem level reliability specs; flow requirements down to silicon, platform, and facilities teams.

  • Lead FMEA, root-cause analysis, and pod fleet failure-data analytics to drive corrective actions and spec updates.

  • Architect RAS features (ECC, memory mirroring, predictive failure, telemetry) and graceful degradation/redundancy against pod-level specs.

  • Partner with facilities on pod power/cooling redundancy (N+1, 2N), thermal margins, and disaster-recovery readiness.

  • Establish HALT/HASS, burn-in, qualification processes; track field returns and KPIs against pod spec.

Qualifications:

Minimum Qualifications:

  • BS/MS/PhD in EE/ME Reliability or related; and/or at least 4-6 yrs experience.

  • Experience authoring and owning reliability specs and requirement flow-down.

  • Strong RAS, FMEA, statistical reliability (Weibull, FIT) skills.

  • Experience with large-scale fleet telemetry and thermal/power redundancy.

Preferred Qualifications:

  • AI cluster operations, data analytics (Python/SQL).

          

Job Type:

Experienced Hire

Shift:

Shift 1 (United States of America)

Primary Location: 

US, Massachusetts, Beaver Brook

Additional Locations:

US, California, Santa Clara

Posting Statement:

All qualified applicants will receive consideration for employment without regard to race, color, religion, religious creed, sex, national origin, ancestry, age, physical or mental disability, medical condition, genetic information, military and veteran status, marital status, pregnancy, gender, gender expression, gender identity, sexual orientation, or any other characteristic protected by local law, regulation, or ordinance.

Position of Trust

N/A

Benefits

We offer a total compensation package that ranks among the best in the industry. It consists of competitive pay, stock bonuses, and benefit programs which include health, retirement, and vacation. Find out more about the benefits of working at Intel.

 

 

Annual Salary Range for jobs which could be performed in the US: $122,440.00-232,190.00 USD

 

 

The range displayed on this job posting reflects the minimum and maximum target compensation for the position across all US locations. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific compensation range for your preferred location during the hiring process.

 

 

Work Model for this Role

This role will require an on-site presence. * Job posting details (such as work model, location or time type) are subject to change.

*

ADDITIONAL INFORMATION: Intel is committed to Responsible Business Alliance (RBA) compliance and ethical hiring practices. We do not charge any fees during our hiring process. Candidates should never be required to pay recruitment fees, medical examination fees, or any other charges as a condition of employment. If you are asked to pay any fees during our hiring process, please report this immediately to your recruiter.
About this role

Summary

Define and maintain reliability specs for AI hardware in large-scale data centers.

Job title

Reliability Engineer

Experience level

4-6 yrs

Minimum experience

4+ years exp

Industry

technology

Location requirements

On-site in Massachusetts or California, USA; no remote work allowed.

Salary

$122k–$232k

Management role

No

Skills & keywords

Required skills

reliability specificationsRASFMEAstatistical reliabilitytelemetry

Preferred skills

AI cluster operationsdata analytics

Specializations

reliabilityRASFMEAthermalpower redundancy
Locations

Structured locations inferred from the posting.

Beaver Brook, Worcester, MA 01602, USA

On-site City

Santa Clara, CA, USA

On-site City