AIML - Machine Learning Researcher - Multimodal Agent

Santa Clara Seattle Until 9/22/2026 First posted July 24, 2026 Last posted July 24, 2026
Job description

The AIML Multimodal Foundation Model Team is pioneering next-generation intelligent agent technologies that combine multimodal reasoning, tool-use, and visual understanding. Our innovative features redefine how hundreds of millions of people utilize their computers and mobile devices for search and information retrieval. Our universal search engine powers search capabilities across a range of Apple products, including Siri, Spotlight, Safari, Messages, and Lookup. Additionally, we develop cutting-edge generative AI technologies based on multimodal large language models to enable innovative features in both Apple’s devices and cloud-based services. As a member of this team, you will design new architectures for multimodal agents, explore advanced training paradigms, and build robust agentic capabilities such as planning, grounding, tool-use, and autonomous task execution. You will collaborate closely with researchers and engineers to bring cutting-edge agent research into production, transforming Apple devices into intelligent partners that help users get things done.

Description

As a member of our fast-paced group, you’ll have the unique and rewarding opportunity to shape upcoming products from Apple. We are looking for people with excellent applied machine learning, computer vision, multimodal LLM, and agent training experience and solid engineering skills.

This role will have the following responsibilities:
- Developing state-of-the-art multimodal foundation models for Apple Intelligence.
- Developing various agent capabilities for multimodal LLMs, including computer use agents, visual tool use, thinking with images, and multimodal web search.
- Developing, fine-tuning, and evaluating domain specific foundation models for various tasks and applications in Apple’s AI powered products
- Conducting applied research to transfer the pioneering research in generative AI to production ready technologies
- Understanding product requirements, translate them into modeling tasks and engineering tasks

Minimum Qualifications

PhD, MS or equivalent experience
Experience in machine learning, deep learning, computer vision, or natural language processing
Proficiency in one of following languages: Python, Go, Java, C++

Preferred Qualifications

Excellent data analytical skills
Good interpersonal skills and team player
PhD preferred

About this role

Summary

Research and develop multimodal foundation models and agent capabilities for Apple.

Job title

AIML - Machine Learning Researcher - Multimodal Agent

Experience level

null

Industry

software

Location requirements

Candidates in Santa Clara or Seattle; remote work not specified.

Salary

Not specified

Management role

No

Skills & keywords

Required skills

PythonGoJavaC++

Preferred skills

data analysisinterpersonal skills

Specializations

machine learningcomputer visionnatural language processingmultimodal modelsagent training
Locations

Structured locations inferred from the posting.

Santa Clara, CA, USA

Work arrangement unknown City

Seattle, WA, USA

Work arrangement unknown City