Senior Research Engineer, Training Data Infrastructure in Foundation Models

Cupertino Until 9/22/2026 4+ years exp First posted July 24, 2026 Last posted July 24, 2026
Job description

We build frontier foundation models that power intelligent experiences at Apple. Our team works across the full training lifecycle: including pre-training foundation models, and developing mid-training approaches that bridge general capability and task-specific performance. What makes our work distinct is that we're engineering models specifically for Apple silicon and optimized for experiences that are private, personal, and deeply integrated into the OS. We're solving frontier problems in reward modeling to resist reward hacking, handling sparse and delayed rewards in agentic settings, and aligning models reliably across the spectrum from open-ended creative tasks to precise, action-taking workflows. If you're drawn to hard problems where the research and the product are inseparable, this is the team.

Description

This position operates at the convergence of Software Engineering and Machine Learning Research. Unlike traditional backend roles, this position requires you to design systems where the outcome is the statistical distribution and quality of data itself. You will work alongside Research Scientists to transform theoretical observations into concrete, scalable engineering solutions. Your core focus will be the architecture of our Data Acquisition, Processing, and Repository Management systems for Large Model training. You will lead technical efforts to enable active, quality-driven data curation, including filtering, deduping, synthetic data generation and data mixing, ensuring our models are trained on the highest-quality information available.

Minimum Qualifications

Education: Bachelor’s degree in Computer Science, Electrical Engineering, or Mathematics.
Technical Expertise: 4+ years of software engineering experience with a specific focus on Data Infrastructure, Distributed Systems, or AI/ML Engineering.
Language Proficiency: Expert fluency in Python, and strong competence in system languages such as C++.
Cloud Architecture: Extensive experience architecting solutions on major public cloud platforms (e.g. GCP) to build scalable data systems (e.g. with Apache Beam, GCS)
Performance Engineering: Deep experience profiling and optimizing high-throughput data systems. Demonstrated ability to debug distributed bottlenecks (e.g., stragglers, I/O saturation), optimize data formats and provide efficient data storage solutions.

Preferred Qualifications

Research Collaboration: Experience working within or closely with ML research organizations (e.g., as a Research Engineer), with an ability to translate research results into engineering implementations.
Domain Knowledge: Familiarity with lifecycle of modern LLM training, end-to-end workflows, and underlying system architecture.
Complex Data Types: Experience in processing complex data modalities beyond plain text, such as source code repositories, images, videos, and audios.

About this role

Summary

Designs scalable data systems for training large models, focusing on data quality and processing.

Job title

Senior Research Engineer, Training Data Infrastructure in Foundation Models

Experience level

4+ years

Minimum experience

4+ years exp

Industry

technology

Location requirements

Remote work allowed; based in Cupertino

Salary

Not specified

Management role

No

Skills & keywords

Required skills

PythonC++distributed systemscloud architectureperformance profiling

Preferred skills

research collaborationml workflowsdata modalities

Specializations

data infrastructuredistributed systemsmachine learningcloud architectureperformance engineering
Locations

Structured locations inferred from the posting.

Cupertino, CA, USA

Remote City