Site Reliability Engineer — Insight Team

Shanghai Until 9/21/2026 3+ years exp First posted July 23, 2026 Last posted July 23, 2026
Job description

The Insight team runs one of Apple's most critical Big Data ecosystems — an exabyte-scale, highly-available infrastructure that underpins manufacturing operations for every Apple product, globally. Every iPhone, iPad, and Mac has touched our systems. We advance technology by relying on each other's strengths and skills to build something bigger than ourselves. For this reason, team culture is central to our values. We value social skills and integrity as much as technical craft. We are looking for extraordinary DevOps with experience building large-scale data platforms, analytic tools and solutions which can help take our platform to the next level. Do you excel in a high-demand setting and exceed expectations, in an environment that requires time-management? The right person will prioritize tasks and complete assignments ahead of schedule. While being a great standout colleague, you will also work independently.

Description

The primary mission of this role is to ensure the stability and performance of our production environment, including the secure and reliable execution of all deployment and monitoring processes. We expect this role to go beyond operations by fully embracing a dev-and-ops mindset. Beyond standard support and troubleshooting, you will leverage AI tools and automation to drive efficiency, applying your engineering expertise to continuously upgrade and optimize our Big Data and microservices ecosystem.

Minimum Qualifications

BS or MS in Computer Science, Software Engineering, or an equivalent technical discipline
Good programming skills in Java, Python or Go
Hands-on experience in SRE, DevOps, or platform engineering, with demonstrated ability to work on technical initiatives
Experience applying AI/ML techniques to development & operations
Proficiency in English for clear technical communication, documentation, and global collaboration
Availability to participate in SRE on-call rotations during China morning hours (8:00 AM CST during US Daylight Saving Time, and 9:00 AM CST during US Standard Time)

Preferred Qualifications

Cloud Technologies: Hands-on experience with cloud platforms (AWS or GCP), containerization and orchestration (Docker, Kubernetes), and multi-region infrastructure design including data residency considerations and IAM.
Distributed Systems & Big Data Platforms: Strong background in microservices, APIs, and messaging/streaming (Kafka, Solace, Pub/Sub), combined with production experience in relational databases (MySQL, PostgreSQL) and big data/storage technologies (Redis, Bigtable, Druid, Elasticsearch, ClickHouse).
CI/CD & Observability: Proven ability to build and maintain CI/CD pipelines or GitOps workflows (ArgoCD, Jenkins, GitHub Actions), and operate observability systems (Grafana, Prometheus, Kibana) including service instrumentation, dashboard design, and alert tuning.
AI tooling & Problem Solving: Forward-looking experience in developing MCP and AI Agentic flows, paired with a flexible and creative approach to solving complex technical problems.

About this role

Summary

Ensure stability and performance of Apple’s big data ecosystems using DevOps, AI, and automation.

Job title

Site Reliability Engineer — Insight Team

Experience level

3+ years

Minimum experience

3+ years exp

Industry

software

Location requirements

Shanghai, on-site preferred, flexibility possible

Salary

Not specified

Management role

No

Skills & keywords

Required skills

JavaPythonDevOpsAI/MLplatform engineering

Preferred skills

AWSGCPDockerKubernetesKafkarelational databasesCI/CDGrafanaPrometheusKibanaproblem solving

Specializations

DevOpsbig datamicroservicescloud platformsAI/ML
Locations

Structured locations inferred from the posting.

Shanghai, China

On-site City