AI Data & Infrastructure Engineer
Apple Sunnyvale, California, United States
Computers and Electronics Manufacturing · 10,001+ employees
About the role
You will lead the end-to-end design, build, and operation of scalable AI data systems and infrastructure to support enterprise GenAI and Agentic AI capabilities. This includes developing data pipelines, implementing RAG architectures, and collaborating with cross-functional teams to deliver production-ready manufacturing solutions.
What they look for
Requirements
Candidates must hold a Bachelor's or Master's degree in a relevant field and possess over 8 years of experience in building scalable data platforms and distributed systems. Proficiency in Python, SQL, and modern data engineering tools like Spark, Kafka, and cloud-native environments is required.
Full description
Imagine what you could do here. At Apple, we believe new insights have a way of becoming excellent products, services, and customer experiences very quickly. Bring passion and dedication to your job and there’s no telling what you could accomplish.
The people here at Apple don’t just build products — they build the kind of wonder that’s revolutionized entire industries. It’s the diversity of those people and their ideas that inspires the innovation that runs through everything we do, from amazing technology to industry-leading environmental efforts. Join Apple, and help us leave the world better than we found it. Manufacturing Systems and Infrastructure (MSI) team is an engineering organization under the Product Operations org. MSI is responsible for the design, development, and maintenance of systems tools, services, and applications required to efficiently run manufacturing operations at scale across global factory sites.
As an AI Data & Infrastructure Engineer with the MSI team, you will won the end to end design, build, and operation of scalable AI data systems that power enterprise GenAI and Agentic AI capabilities. Your work spans core platform services, data pipeline development and infrastructure provisioning, enabling manufacturing workflow automation through Agents and Agent skills.
Description
Create robust, scalable architectures for systems that handle data orchestration for AI features
Design, build, and maintain scalable AI data platforms, services, and APIs that support and enable Agentic AI workflow development & automation.
Develop data ingestion, transformation, and publishing pipelines for structured, unstructured, and multimodal data.
Design and implement Retrieval-Augmented Generation (RAG) pipelines, embedding workflows, vector database integrations, and metadata services for enterprise AI applications.
Be able to quickly build an idea so you and the team can work with hands-on products. Then iterate on the best of those prototypes.
Build and integrate tools that help make complex AI systems observable, understandable and debuggable
Strong understanding of distributed systems, parallel computing, and performance optimization
Collaborate with AI/ML engineers, software engineers, product teams, and domain experts to define AI data requirements and deliver production-ready data solutions.
Optimize platform scalability, reliability, performance, security, and cost across cloud-native environments and Agentic systems.
Ability to clearly communicate complex technical problems and collaborate with partners to develop solutions
Minimum Qualifications
Bachelor's or Master's degree in Computer Science, Software Engineering, Data Engineering, or a related field. 8+ Experience designing and building scalable data platforms and distributed systems. Strong programming skills in Python and SQL, with proficiency in Java or Scala preferred. Experience with Airflow, Kubeflow, or MLflow to build and orchestrate scalable AI data pipelines. Experience building scalable batch and streaming data pipelines using Spark (PySpark), Kafka, Airflow, and Ray, with proficiency in Pandas and modern data lake/lakehouse architectures (e.g., Iceberg, Delta Lake). Experience using modern development tools, including AI-assisted coding tools, while applying sound engineering judgment to review, validate, and improve generated code. Knowledge of RAG architectures, embedding generation, vector databases, and AI data preparation for LLMs and agentic AI. Experience with cloud platforms (AWS, Azure, or GCP), Kubernetes, Docker, CI/CD, and Infrastructure as Code. Demonstrated ability to understand complex user workflows, translate them into practical technical solutions, and collaborate across teams to deliver measurable outcomes. Excellent communication and collaboration skills, with the ability to translate technical concepts into clear, business focused insights.
Preferred Qualifications
Experience with or a strong understanding of Generative AI, LLMs, Agentic Systems, or RAG (Retrieval-Augmented Generation) workflows. Experience integrating LLMs into existing systems Experience with API design, both for other engineers to use, but also for AI systems. Experience working with manufacturing, operational, IoT, or industrial data platforms. Demonstrated ability to lead technical initiatives and mentor engineers.