ASE Compute - Site Reliability Engineering (SRE) Manager

Seattle Until 9/22/2026 5+ years exp First posted July 24, 2026 Last posted July 24, 2026
Job description

Apple's ASE Compute team builds and operates the private cloud infrastructure that powers Apple services at massive scale. Our platform delivers bare-metal Kubernetes clusters and virtualized environments to thousands of engineers across the company. We are looking for an SRE Manager to lead a team that keeps this infrastructure reliable, performant, and ready for the next order-of-magnitude growth.

Description

This is a hands-on leadership role. You will set the technical direction for reliability and operational excellence while mentoring engineers, driving automation, and partnering closely with software and infrastructure teams to ship improvements that matter.

You will have direct impact on the platform that underpins Apple's most critical services. You will work alongside world-class engineers solving problems at a scale few organizations encounter — and you will build a team culture that makes reliability engineering sustainable and rewarding.

Minimum Qualifications

5+ years of engineering management experience leading infrastructure or SRE teams
Deep experience operating large-scale, multi-tenant Kubernetes environments in production
Strong systems background — comfortable troubleshooting across the full stack (network, OS, container runtime, application)
Experience with configuration management at scale (Puppet, Ansible, or equivalent)
Track record of building high-performing teams through coaching, clear expectations, and psychological safety
Demonstrated ability to drive cross-functional initiatives to completion
Strong written and verbal communication skills

Preferred Qualifications

Experience with third-party cloud platforms (AWS, GCP, or Azure)
Familiarity with bare-metal provisioning and lifecycle management at datacenter scale
Experience with Java, Go, or Python services in production
Understanding of cloud-native observability (Prometheus, Thanos, Splunk, or similar)
CNCF Certified Kubernetes Administrator (CKA) or equivalent hands-on certification
Experience running infrastructure as an internal managed service with defined SLAs

About this role

Summary

Lead a team to ensure reliable, scalable cloud infrastructure and automate operations.

Job title

ASE Compute - Site Reliability Engineering (SRE) Manager

Experience level

5+ years

Minimum experience

5+ years exp

Industry

technology

Location requirements

Seattle, on-site with remote work possible

Salary

Not specified

Management role

Yes

Skills & keywords

Required skills

infrastructure managementkubernetesconfiguration managementcommunication

Preferred skills

awsgcpazurebare-metal provisioningpythongoprometheus

Specializations

kubernetesinfrastructurereliability engineeringautomationcloud platforms
Locations

Structured locations inferred from the posting.

Seattle, WA, USA

Hybrid City