Machine Learning Engineer - Speech & Multimodal Language Modeling

Cupertino Until 9/22/2026 2+ years exp First posted July 24, 2026 Last posted July 24, 2026
Job description

Apple is where individual imaginations gather together, committing to the values that lead to great work. Every new product we build, service we create, or experience we deliver is the result of us making each other’s ideas stronger. The diversity of our people and their thinking inspires the innovation that runs through everything we do. When we bring everybody in, we can do the best work of our lives. Here, you’ll do more than join something — you’ll add something.

Description

The Special Projects team at Apple is developing novel user-facing features that leverage the
multimodal capabilities of state-of-the-art foundation language models. We are looking for a
highly skilled Machine Learning Engineer to build and evaluate these experiences, with a
specific focus on Multimodal and Speech Language Models. A successful candidate is
experienced in evaluating complex foundation model-driven systems end-to-end, translating
subjective product requirements into objective criteria, has strong statistical analysis skills, and
has worked with Speech Language Models.

Minimum Qualifications

Master’s degree in Computer Science or Machine Learning
2+ years of hands-on experience building and evaluating generative AI models
Proficiency in Python and ML frameworks (Pytorch or Tensorflow)

Preferred Qualifications

PhD in Computer Science, Machine Learning, Statistics, or other STEM field
5+ years of hands-on experience with SpeechLMs or LLMs
Experience with large-scale audio data processing on distributed systems
Experience with prompt evaluation and optimization for generative AI models
Proficiency in training, fine-tuning, and evaluation of foundation models and frameworks
A track record of publications or technical presentations in Machine Learning journals or
conferences
Excellent communication skills and cross-functional collaboration

About this role

Summary

Develop and evaluate multimodal and speech language models; analyze complex AI systems.

Job title

Machine Learning Engineer - Speech & Multimodal Language Modeling

Experience level

2+ years

Minimum experience

2+ years exp

Industry

technology

Location requirements

Must be located in Cupertino, remote not specified.

Salary

Not specified

Management role

No

Skills & keywords

Required skills

pythonpytorchtensorflowevaluating ai models

Preferred skills

phd in stemspeechlmslarge-scale audio processingprompt evaluationfine-tuningpublications

Specializations

multimodalspeech language modelsgenerative aifoundation models
Locations

Structured locations inferred from the posting.

Cupertino, CA, USA

On-site City