Research Scientist – World Modeling, Data
Institute of Foundation Models Sunnyvale, California, United States · $150K–$400K/yr
Research Services · 51-200 employees
About the role
You will design and build scalable data pipelines for curating and annotating billion-scale visual data to improve world model capabilities. You will also collaborate with engineering teams to deploy these systems on AWS infrastructure and in-house clusters.
What they look for
Requirements
The role requires a PhD in Machine Learning or Computer Science, or equivalent industry experience. Candidates should possess strong expertise in deep learning frameworks and experience with large-scale data curation and annotation.
Benefits
Full description
About the Institute of Foundation Models
We are a dedicated research lab for building, understanding, using, and risk-managing foundation models. Our mandate is to advance research, nurture the next generation of AI builders, and drive transformative contributions to a knowledge-driven economy.
As part of our team, you’ll have the opportunity to work on the core of cutting-edge foundation model training, alongside world-class researchers, data scientists, and engineers, tackling the most fundamental and impactful challenges in AI development. You will participate in the development of groundbreaking AI solutions that have the potential to reshape entire industries. Strategic and innovative problem-solving skills will be instrumental in establishing MBZUAI as a global hub for high-performance computing in deep learning, driving impactful discoveries that inspire the next generation of AI pioneers.
Position Summary
We are looking for Research Scientists to tackle the data problems that determine our world model capabilities. You will design, develop, and run state-of-the-art systems for curating, ingesting, filtering, annotating, and training with billion-scale visual data. You will work on problems including, but not limited to, motion scoring, camera control, and video understanding. You will translate our team’s human efforts into durable automated systems, whose performance is quantitatively verifiable.
\n
\n$150,000 - $400,000 a year
Key Responsibilities
• Translate requirements from downstream training, application and eval teams to actionable data work.
• Design and build scalable data pipelines.
• Collaborate with engineering to scale data pipelines on AWS infrastructure and in-house clusters.
• Develop state-of-the-art visual understanding systems to accurately label and annotate images and videos, such as for detecting AI-generated content.
Academic Qualifications
• PhD in Machine Learning or Computer Science, or equivalent industry experience.
Professional Experience
Necessary:
• Strong communication and collaboration skills for effective cross-functional teamwork.
• Ability to navigate ambiguity and drive projects in rapidly evolving research areas.
• Exceptional problem-solving and troubleshooting skills to tackle complex technical challenges.
Preferred:
• Experience in building and optimizing large-scale video data pipelines.
• Experience with large scale video curation.
• Experience with data filtering, particularly AI-generated content detection and duplicate detection.
• Experience with image or video data annotation.
• Experience with video data curriculum.
• Strong systems and engineering expertise in deep learning frameworks such as PyTorch.
• Research contributions to top-tier conferences or journals (e.g., ICML, ICLR, NeurIPS, ACL, CVPR, COLM, etc.), with published work in relevant domains.
\nVisa Sponsorship
This position is eligible for visa sponsorship.
Benefits Include
*Comprehensive medical, dental, and vision benefits
*Bonus
*401K Plan
*Generous paid time off, sick leave and holidays
*Paid Parental Leave
*Employee Assistance Program
*Life insurance and disability