Apple

Staff Machine Learning Engineer, Siri Attention and Invocation

Apple Zurich, Zurich, Switzerland

Computers and Electronics Manufacturing · 10,001+ employees

5 h ago
machine-learning Principal (10+ yrs) Full-time Switzerland
Log in to apply, save this posting, or score it against your profile with AI.

About the role

Lead the technical vision for generating realistic, expressive synthetic speech and visual representations for conversational agents. Drive the development of audio and video generation capabilities while mentoring senior and junior engineers across the team.

What they look for

Generative AI Machine Learning Speech Synthesis Computer Vision Diffusion Models Autoregressive Models GANs Python PyTorch TensorFlow Distributed Training Multimodal Modeling Audio Processing Video Generation Technical Leadership Software Engineering

Requirements

Requires a Master’s or PhD in Computer Science, Electrical Engineering, Machine Learning, or a related field. Candidates must have deep hands-on experience with generative audio and video architectures and a proven track record of leading complex projects from research to production.

Full description

As part of Siri Attention and Invocation, we collaborate to deliver the next revolution in human-computer interaction, to inspire and create groundbreaking technology for large-scale systems spanning speech, vision, and generative AI to overcome real-world challenges through innovation and user-centered design that improves the daily experience of millions of our customers.

Description

We are seeking an exceptional Staff Machine Learning Engineer to lead the development of audio and video generation capabilities that bring conversational agents to life. In this role, you will drive the technical vision for generating realistic, expressive synthetic speech and visual representations, mentor senior and junior engineers, and shape the roadmap for multimodal generative experiences across our products.

Minimum Qualifications

Master’s or PhD in Computer Science, Electrical Engineering, Machine Learning, or a related field, or equivalent practical experience Deep hands-on experience with generative audio and/or video architectures (e.g., diffusion models, autoregressive models, GANs, VAEs, end-to-end neural synthesis) Demonstrated ability to lead complex, ambiguous projects from research through production, and to make sound technical tradeoffs under real-world constraints (quality, latency, compute) Experience evaluating generative model outputs, including both objective metrics and perceptual/subjective quality assessment

Preferred Qualifications

Proven experience building and shipping machine learning systems in production, with significant focus on generative modeling Excellent collaboration and communication skills, with a track record of working across research, engineering, and product teams Strong software engineering skills, with experience designing scalable ML systems and pipelines (e.g., Python, PyTorch/TensorFlow, distributed training infrastructure) Experience with speech synthesis (TTS), voice conversion, audio acoustics/background modeling, or conversational AI systems Experience with generative video/animation techniques (e.g., facial animation, lip-sync, avatar rendering, video diffusion) Publications in generative modeling, speech, audio, or computer vision at top-tier venues (e.g., NeurIPS, ICML, ICASSP, CVPR, Interspeech) Experience deploying real-time or low-latency generative models at scale Familiarity with multimodal modeling (joint audio-visual generation, cross-modal conditioning) Prior experience mentoring engineers or leading technical direction for a team

Similar roles