Machine Learning Engineer, Foundation Model Services
Apple Seattle, Washington, United States
Computers and Electronics Manufacturing · 10,001+ employees
About the role
You will build production-grade solutions to serve foundation models in real-time and collaborate with researchers to develop inference for cutting-edge architectures. Additionally, you will create tools to identify and resolve performance bottlenecks across hardware and various use cases.
What they look for
Requirements
Candidates must have at least 5 years of industry experience in ML technologies and 2+ years in building production software. A bachelor's degree in Computer Science or a related field is required, along with proficiency in Python or Go and experience with cloud infrastructure.
Full description
Do you think differently? Are you eager to break the status quo, bold and ambitious, unafraid to take risks, and passionate about building best-in-class technology? If so, there's no better place to do it than Apple. The Foundation Model Services team builds the frameworks, services, and tools that run Apple's largest foundation models in production. Our infrastructure powers intelligent experiences across products people use every day — Search, Music, TV, the App Store, Messages, Photos, Spotlight, Safari, Siri, and more — serving millions of queries at incredibly low latency while drawing every ounce of performance from our hardware. Join us and you'll help bring intelligence to billions of users around the world, working on optimizing and serving large language, vision, and speech models at Apple's scale.
Description
Work closely with product teams to build production-grade solutions that launch models serving customers in real time. Partner with foundation model researchers to prototype and develop inference for cutting-edge model architectures, and build tools that help us understand and remove performance bottlenecks across different hardware and use cases. Write high-quality code, learn quickly in a fast-moving field, and grow your impact as you take on larger pieces of the system.
Minimum Qualifications
2+ years of industry experience building and shipping production software and/or machine learning systems. Proficiency in a modern programming language such as Go or Python. Experience deploying and operating services on a cloud platform (AWS, Azure, GCP, or equivalent) using containers and Kubernetes/Docker. 5 year+ industry experience in ML technologies (LLMs, Machine Learning, NLP, Information Retrieval, Statistics). Experience building or operating high-throughput, low-latency services. Strong communication and collaboration skills, with the ability to partner across research and product teams. Bachelor’s degree or higher in Computer Science or related technical field.
Preferred Qualifications
Familiarity with Nvidia TensorRT-LLM, vLLLM, DeepSpeed, Nvidia Triton Server etc.
Similar roles
-
Senior Staff Machine Learning Engineer, Feed Relevance
Reddit United States · $266K–$372K/yr
-
Machine Learning Engineer, II - 3D Perception
Torc Robotics Ann Arbor, Michigan, United States · $153K–$184K/yr
-
Machine Learning Operations Engineer
4MindsAI Inc. Dallas, Texas, United States · $130K–$200K/yr
-
Machine Learning Engineer II/III (Applied Research & Model Development)
PathAI Boston, Massachusetts, United States · $107K–$200K/yr
-
Senior Machine Learning Engineer Mandarin Speaker
AHU Technologies Inc United States · $150K–$300K/yr
-
Machine Learning Programmer, Memory
Epic Games Cary, North Carolina, United States · $184K–$271K/yr