Data Platform Infra Site Reliability Engineer

London Until 9/27/2026 First posted July 29, 2026 Last posted July 29, 2026
Job description

At Apple, we believe that innovation flourishes in an environment where ideas are challenged, collaboration is encouraged, and technology is pushed to its limits. This environment is only possible when diverse minds come together, bringing unique perspectives and experiences. Our people and their ideas inspire innovation in everything we do. Imagine what you could accomplish here! Join Apple and help us make the world a better place. As an SRE on our team, you'll own the reliability, performance, and scale of the distributed storage and data platform systems that power Apple's services. You'll debug replication and consensus failures, tune systems at petabyte scale, and write production code that operates the platform. We firmly believe in ownership, with software engineers accountable for the code they write.

Description

The Apple Services Engineering (ASE) organisation builds and provides systems and infrastructure that fuel Apple’s services — iCloud, iTunes, Siri, and Maps. Our team builds and operates the data platform infrastructure behind them, keeping petabyte-scale workloads fast, resilient, and reliable.

The platform runs on large-scale distributed systems, including object stores, databases, and data pipelines, on Linux across private and hybrid cloud. You'll work on storage engines, distributed consensus, and data-flow internals, partnering with development teams on system-wide architecture rather than individual components.

Minimum Qualifications

Experience in managing and scaling large-scale distributed systems in a private or hybrid cloud environment.
Comfortable designing, writing, and releasing production code in languages such as Go or Python.
Able to debug and reason about how distributed systems fail and perform at scale.
Willingness to take part in on-call rotations and incident response to keep critical systems healthy.

Preferred Qualifications

Contributions to distributed-systems internals, open-source data infrastructure, or storage/database engines.
Experience defining SLIs/SLOs, building observability, and using error budgets to drive reliability decisions.
A good grasp of Unix internals and networking fundamentals.
Experience with data migration, disaster recovery, or capacity planning at scale.

About this role

Summary

Manage and optimize large-scale distributed storage and data platform systems.

Job title

Data Platform Infra Site Reliability Engineer

Experience level

none

Industry

technology

Location requirements

London, remote work not specified

Salary

Not specified

Management role

No

Skills & keywords

Required skills

managing large-scale distributed systemsGoPythondebug distributed systemson-call incident response

Preferred skills

distributed-systems internalsopen-source data infrastructurestorage enginesSLIs/SLOsUnix internalsnetworkingdata migrationdisaster recoverycapacity planning

Specializations

distributed systemsstorage enginesclouddatabases
Locations

Structured locations inferred from the posting.

London, UK

Work arrangement unknown City