Apple

AIML - Quality Engineering Manager, Responsible AI and Safety

Apple Cupertino, California, United States

Computers and Electronics Manufacturing · 10,001+ employees

4 h ago
engineering-manager Senior (5-10 yrs) Full-time United States
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

Lead the Quality Platform Engineering function to provide fast, trustworthy safety signals for Apple Intelligence and extend testing coverage across hardware platforms. Build and manage a high-performing team while establishing a flywheel that connects engineering insights to product evaluations and leadership decision-making.

What they look for

Machine Learning AI Safety Quality Engineering Python Test Infrastructure CI/CD Evaluation Pipelines Agentic Systems Product Management Team Leadership Data Analysis Hardware Testing Generative Models Responsible AI Cross-platform Testing

Requirements

Requires 5+ years of technical leadership experience with a strong background in ML evaluation, production systems, and test infrastructure. Candidates must possess excellent communication skills, the ability to make product recommendations in ambiguous environments, and proficiency in Python.

Full description

Apple's Responsible AI and Safety team focuses on innovative technologies, methodologies, and research to enable fantastic user experiences and to push the frontier of machine learning. Our team is looking to hire a leader with a strong track record in Applied Research, who is passionate about ML and foundation models with a focus on responsibility, fairness, and safety. In this role, you will lead the research and application of ML methods for technologies that power breakthrough user experiences while upholding Apple's values, privacy, and quality standards.

Description

This role leads Apple's Quality Platform Engineering function within Responsible AI — the team responsible for giving engineering teams across Apple Intelligence fast, trustworthy safety signal before code merges, and for extending safety testing coverage to every hardware platform and form factor Apple ships. Most of this role's impact will come from two things: building a strong team and a healthy team culture from the ground up, and establishing the flywheel that connects Quality Platform Engineering to Product Evaluations & Research and Post-Ship Insights — so fast signal, coverage gaps, and pipeline improvements consistently translate into real engineering decisions and clear leadership visibility, rather than one-off dashboards nobody acts on.You should be technically fluent — comfortable with evaluation pipelines, production ML and agentic systems, and cross-platform testing — enough to earn credibility with the team, ask sharp questions, and evaluate hard tradeoffs with good judgment. This role moves at a fast pace, and you should be comfortable making product recommendations in ambiguous situations, often with limited or imperfect data. Just as important are excellent communication, strong product sense to prioritize a team's limited capacity against Apple's highest-risk platforms and use cases, and a track record of hiring and developing technical teams.Prior experience managing or building technical teams is required. Prior exposure to ML evaluation, test infrastructure, or safety-adjacent engineering work is strongly preferred.

Minimum Qualifications

5+ years of technical team management or leadership experience Experience with ML evaluation, production ML systems, or test/CI infrastructure at scale Strong engineering skills and experience writing production-quality code (Python or similar) Experience working across multiple platforms or hardware form factors, or a demonstrated ability to ramp quickly across unfamiliar platforms Experience working with human-labeled or crowd-sourced evaluation data, including reasoning about label noise and inter-rater agreement

Preferred Qualifications

Experience working on Responsible AI, AI safety, or trust & safety-adjacent engineering Experience with generative model evaluation and common failure modesStrong organizational and operational skills working with large, multi-functional, diverse teams MS or PhD in Computer Science, Machine Learning, Statistics, or related field, or equivalent experience Familiarity with hardware/platform-specific testing considerations (e.g., on-device constraints, new form factors)

Similar roles