Apple

Machine Learning Engineer - Multimodal Intelligence

Apple Sunnyvale, California, United States

Computers and Electronics Manufacturing · 10,001+ employees

4 h ago
machine-learning Mid (2-5 yrs) Full-time United States
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

You will build and scale data curation and training pipelines to turn multimodal foundation models into production-ready Apple Intelligence features. This role involves optimizing models for on-device execution while ensuring robust performance under strict latency, memory, and power constraints.

What they look for

Deep learning Multimodal systems Python PyTorch JAX Computer vision Data curation Model training Fine-tuning Model evaluation On-device deployment Distributed systems Latency optimization Memory optimization Research prototyping Production systems

Requirements

Candidates must have at least 3 years of industry experience in deep learning and proficiency in Python with PyTorch or JAX. A degree in Computer Science or a related field is required, with a preference for advanced degrees and experience in multimodal foundation models.

Full description

Imagine what you could do here. At Apple, new ideas have a way of becoming extraordinary products, services, and customer experiences very quickly. Bring passion and dedication to your job and there's no telling what you could accomplish. Multifaceted, amazing people and inspiring, innovative technologies are the norm here. The people who work here have reinvented entire industries with all Apple Hardware products. The same passion for innovation that goes into our products also applies to our practices, strengthening our commitment to leave the world better than we found it. Join us in this truly exciting era of Artificial Intelligence to help deliver the next groundbreaking Apple products and experiences!

The Multimodal Intelligence team builds and ships the Computer Vision and Machine Learning systems behind Apple Intelligence — spanning data collection and curation, training and fine-tuning, evaluation, optimization, and on-device deployment. Our team has an established track record of delivering features that combine Apple's sensing hardware with large foundation models, including Visual Intelligence and the on-device foundation models that power text and visual understanding across iPhone, iPad, Mac, and Apple Vision Pro. We are focused on building experiences where a device can see, read, and reason about the world around it — privately, responsively, and on-device wherever possible.

Description

We are looking for a Machine Learning Engineer to build the pipelines, infrastructure, and production systems that turn multimodal foundation models into shipping Apple Intelligence features. You will own end-to-end model delivery: building and scaling data curation and training pipelines, fine-tuning and optimizing large multimodal models for on-device and hybrid execution, standing up reproducible evaluation and regression testing for text and visual understanding, and hardening promising approaches into robust, maintainable production systems under real latency, memory, power, and privacy constraints.

You will work closely with modeling, platform, hardware, and product engineering teams across Apple — taking future hardware design and product needs into account as you make implementation decisions — and you will have the opportunity to collaborate broadly to deliver the best possible products.

Minimum Qualifications

Experience in deep learning with demonstrated work in at least one area of multimodal systems (e.g., vision, language, video, audio, etc.) Proficiency in Python and in a modern deep learning framework such as PyTorch or JAX Experience with rapid prototyping, reproduction, and validation of research ideas Ability to work in a collaborative environment Ability to communicate the results of analyses in a clear and effective manner BS and a minimum of 3 years of relevant industry experience

Preferred Qualifications

Master's or PhD, or equivalent practical experience, in Computer Science, Computer Vision, Machine Learning, or related technical field Deep expertise in multimodal foundation models, with a focus on practical applications Track record of translating research into practical applications either through published work or industry experience Strong applied research experience in at least one major area of model development (data curation, pre-training, fine-tuning, alignment, or evaluation), particularly as it applies to multimodal systems Experience with large-scale training pipelines, including working with large datasets and scaling models across distributed systems Experience bridging research ideas with production constraints

Similar roles