Senior Software Engineer, AI/ML System Infrastructure
About the Job
Google's software engineers develop the next-generation technologies that change how billions of users connect, explore, and interact with information and one another. Our products need to handle information at massive scale, and extend well beyond web search. We're looking for engineers who bring fresh ideas from all areas, including information retrieval, distributed computing, large-scale system design, networking and data storage, security, artificial intelligence, natural language processing, UI design and mobile; the list goes on and is growing every day. As a software engineer, you will work on a specific project critical to Google’s needs with opportunities to switch teams and projects as you and our fast-paced business grow and evolve. We need our engineers to be versatile, display leadership qualities and be enthusiastic to take on new problems across the full-stack as we continue to push technology forward.As the Senior Software Engineer, you will drive software development for Tensor Processing Unit (TPU) system control planes. You will design and implement health management systems that rely on hardware telemetry. You will build analytics for detecting hardware problems and develop different health rules specializing in detecting ICI, OCS, EMD, Tross, and out of band entities failure modes. You will build algorithms to generate correlated failures and suggest actions for repair workflows, integrating with both TPU cluster and cloud infrastructure. You will also build the Diagnoser for the specialized health rules to maintain and operate these TPU clusters.
Google Cloud accelerates every organization’s ability to digitally transform its business and industry. We deliver enterprise-grade solutions that leverage Google’s cutting-edge technology, and tools that help developers build more sustainably. Customers in more than 200 countries and territories turn to Google Cloud as their trusted partner to enable growth and solve their most critical business problems.Individual pay is determined by factors including job-related skills, experience, and relevant education or training.US: $174000 - $252000 (USD) + 15% bonus target + equity + benefits
Learn more about benefits at Google.
Responsibilities
- Build and develop systems to integrate in continuous running of TPU AI infrastructure.
- Work cross-functionally to define the requirements for Diagnoser, and how will the customer benefit from it and define the Critical User Journeys (CUJs).
- Influence and align the cross-functional teams on roadmap of PodCare and Diagnoser.
- Deliver high quality code and timely project while working with cross-functional teams.
- Contribute to CI/CD pipeline and integration and regression test for health monitoring system.
Qualifications
Minimum qualifications:
- Bachelor’s degree or equivalent practical experience.
- 5 years of experience working with Go.
- 3 years of experience with developing large-scale infrastructure, distributed systems or networks, or experience with compute technologies, storage or hardware architecture.
- 3 years of experience in distributed computing.
- 3 years of experience in infrastructure design.
- 3 years of experience in system architecture.
Preferred qualifications:
- Master's degree or PhD in Computer Science or related technical field.
- 5 years of experience with data structures and algorithms.
- 5 years of experience with C and C++.
- 1 year of experience in a technical leadership role.
- Experience developing accessible technologies.
Additional Information
In accordance with Washington state law, we are highlighting our comprehensive benefits package, which is available to all eligible US based employees. Benefits for this role include:- Health, dental, vision, life, disability insurance
- Retirement Benefits: 401(k) with company match
- Paid Time Off: 20 days of vacation per year, accruing at a rate of 6.15 hours per pay period for the first five years of employment
- Sick Time: 40 hours/year (increased to 69 hours/year for Seattle) including 5 discretionary sick days per instance
- Maternity Leave (Short-Term Disability + Baby Bonding): 28-30 weeks
- Baby Bonding Leave: 18 weeks
- Holidays: 13 paid days per year
Summary
Develop and maintain TPU system control, health management, and diagnostic systems.
Job title
Senior Software Engineer, AI/ML System Infrastructure
Experience level
5+ years
Minimum experience
3+ years exp
Industry
software
Location requirements
Remote work allowed; based in USA
Salary
$174k–$252k
Management role
No
Required skills
Preferred skills
Specializations
Structured locations inferred from the posting.
United States
Sunnyvale, CA, USA
Kirkland, WA, USA