SOC Quality and Reliability Engineer, Google Cloud

Tel Aviv, Israel Haifa, Israel Until 9/19/2026 8+ years exp First posted July 21, 2026 Last posted July 21, 2026
Job description

About the Job

In this role, you’ll work to shape the future of AI/ML hardware acceleration. You will have an opportunity to drive cutting-edge TPU (Tensor Processing Unit) technology that powers Google's most demanding AI/ML applications. You’ll be part of a team that pushes boundaries, developing custom silicon solutions that power the future of Google's TPU. You'll contribute to the innovation behind products loved by millions worldwide, and leverage your design and verification expertise to verify complex digital designs, with a specific focus on TPU architecture and its integration within AI/ML-driven systems.

Google’s data centers are the most advanced in the world. In this role, you will help build the SoC’s that power these data centers by driving quality and reliability processes from the integrated circuit perspective. You will create silicon and follow it into the field (and back) to drive improvements for the next generations of chips.

You will have an understanding of IC flows, wafer processing, testing, qualification, yield, reliability, and failure analysis. You will work with various cross functional teams to develop quality and reliability specifications, develop and deploy design guidelines, and develop and execute and test plans. Within the larger organization you will collaborate with global hardware quality and reliability teams, silicon design, validation and engineering teams.

The AI and Infrastructure team is redefining what’s possible. We empower Google customers with breakthrough capabilities and insights by delivering AI and Infrastructure at unparalleled scale, efficiency, reliability and velocity. Our customers include Googlers, Google Cloud customers, and billions of Google users worldwide.

We're the driving force behind Google's groundbreaking innovations, empowering the development of our cutting-edge AI models, delivering unparalleled computing power to global services, and providing the essential platforms that enable developers to build the future. From software to hardware our teams are shaping the future of world-leading hyperscale computing, with key teams working on the development of our TPUs, Vertex AI for Google Cloud, Google Global Networking, Data Center operations, systems research, and much more.

Responsibilities

  • Drive the strategic definition and development of design-for-reliability (DfR) guidelines, collaborating with cross-functional subject matter experts to integrate reliability into early design stages.
  • Define and lead the development of qualification hardware and test methodologies, managing internal teams and external vendors to ensure silicon and package verification.
  • Execute comprehensive silicon and package qualification programs (including high-temperature operating life (HTOL), early life failure rate (ELFR), electrostatic discharge/latch-up (ESD/LU), biased highly accelerated stress test (b/HAST), etc.) and conduct failure analysis to resolve quality issues.
  • Extract and analyze data from qualification programs, high-volume manufacturing, and field returns to identify failure mechanisms and trends for yield and reliability optimization.
  • Develop and implement physics-based statistical Quality and Reliability models (e.g., ELF, time-dependent dielectric breakdown (TDDB), negative bias temperature instability (NBTI) to predict device failure mechanisms and lifetime behaviors.

Qualifications

Minimum qualifications:

  • Bachelor's degree in Electrical Engineering, Materials Science, Physics, or a related field or equivalent practical experience.
  • 8 years of experience in IC silicon quality or reliability.
  • Experience leading the product reliability lifecycle from post-tapeout through high-volume manufacturing.
  • Experience with semiconductor complementary metal-oxide-semiconductor (CMOS) technology, device physics, and failure mechanisms.

Preferred qualifications:

  • Master's degree in Electrical Engineering, Materials Science, or related field.
  • Expertise in statistical data analysis using tools such as JMP, Python, or JSL.
  • Knowledge of design-for-reliability (DfR) rules and implementation techniques.
  • Familiarity with electrical failure analysis (EFA) and physical failure analysis (PFA) techniques.
  • Track record with silicon reliability on process nodes and advanced packaging technologies.
About this role

Summary

Lead reliability and quality assurance for Google's AI hardware, focusing on silicon verification and failure analysis.

Job title

SOC Quality and Reliability Engineer, Google Cloud

Experience level

8+ years

Minimum experience

8+ years exp

Industry

technology

Location requirements

On-site in Israel, no remote work specified.

Salary

Not specified

Management role

No

Skills & keywords

Required skills

IC silicon qualityreliabilityfailure analysissemiconductorsqualification testing

Preferred skills

statistical data analysisDfR rulesEFAPFAadvanced packaging

Specializations

IC designreliabilityquality assurancesemiconductorsfailure analysis
Locations

Structured locations inferred from the posting.

Tel Aviv-Yafo, Israel

On-site City

Haifa, Israel

On-site City