Yum!

Manager, Machine Learning

Yum! Ho Chi Minh City, Vietnam

Restaurants · 10,001+ employees

3 h ago
machine-learning Senior (5-10 yrs) Full-time Vietnam
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

The manager will lead a 24/7 distributed AI operations team to ensure system reliability and scalability while overseeing AI execution initiatives in Asian markets. They will partner with cross-functional teams to define operational playbooks, manage incident responses, and translate business needs into executable AI strategies.

What they look for

AI Operations MLOps Team Leadership Incident Management SLA Management Cloud Platforms Kubernetes Docker Python Bash Prometheus Grafana Datadog PagerDuty Machine Learning Lifecycle Management Strategic Planning

Requirements

Candidates must hold a Bachelor's or Master's degree in a technical field and possess at least 5 years of experience in AI/ML operations or DevOps, including 2 years in a leadership role. Strong proficiency in monitoring frameworks, cloud platforms, and the ability to collaborate with global stakeholders is essential.

Benefits

100% salary during probation 18 days annual leave Five recharge days Paid birthday leave Work from home 2 days per week Half-day Fridays Full salary insurance 13th-month bonus Advanced health insurance Team and engagement activities LinkedIn learning and training courses MacBook and monitors

Full description

Our story might surprise you. We’re the world’s largest restaurant company—encompassing KFC, Taco Bell, and Habit Burger & Grill—but there’s a lot more going on behind the scenes than just serving great food. We put this delicious food in the hands of customers through apps, websites, kiosks, POS, and other digital dining experiences. 

We are seeking a strategic and hands-on AI Operations & Asia AI Execution Manager with a dual mandate. Approximately 80% of the initial role will lead the AI Operations and Support function, ensuring the reliability, scalability, and efficiency of AI/ML systems in production. Approximately 20% will focus on AI execution in Asia, initially for KFC in one to two priority markets. The leader will oversee a 24/7 operations team across Vietnam and Colombia, establish playbooks and SLAs, and partner closely with Data Scientists, AI Engineers, and Machine Learning Engineers on support strategies.

In parallel, this leader will partner with KFC business and AI leaders in the selected Asian markets to identify where existing global AI capabilities can create value, where new capabilities are needed, and how opportunities should be translated into executable initiatives. Working closely with the Global AI Leader and regional AI leaders, they will draw on Data Science and Machine Learning Engineering resources in the United States, Colombia, and India. As this model proves value, the role is expected to expand its market scope and regional portfolio ownership and help build local AI delivery capability in Asia, likely based in Vietnam.

Responsibilities

Team Leadership & Management:

  • Build, lead, and manage a distributed 24/7 AI Operations team based in Vietnam and Colombia, responsible for monitoring and addressing AI system issues.
  • Develop staffing plans, schedules, and rotations to ensure continuous coverage, including weekends and holidays.
  • Foster a collaborative and high-performance culture within the team, emphasizing learning, accountability, and continuous improvement.
  • Provide mentorship and professional development opportunities to team members across Vietnam and Colombia, fostering strong collaboration across locations and time zones.

Operational Excellence:

  • Design, implement, and maintain comprehensive operational playbooks for AI system monitoring, troubleshooting, and issue resolution.
  • Set, monitor, and ensure adherence to service-level agreements (SLAs) for issue response and resolution times.
  • Lead incident response efforts, ensuring timely resolution and thorough post-mortems for continuous improvement.

Onboarding Support:

  • Partner with cross-functional teams to participate in the onboarding of new stores, ensuring AI systems, data pipelines, and operational processes are successfully integrated.
  • Lead AI operational responsibilities after the initial onboarding phase, ensuring a seamless transition to steady-state operations.
  • Provide training and documentation to ensure store teams are familiar with AI-supported workflows and issue resolution processes

Collaboration:

  • Work closely with Data Scientists, AI Engineers, and MLEs to understand AI system architectures and define operational support requirements.
  • Coordinate with product and engineering teams to ensure operational feedback informs system design and deployment processes.
  • Collaborate with the AI team to develop strategies for model retraining, scaling, and lifecycle management.

Asia AI Growth & Execution:

  • Partner initially with KFC business leaders and regional AI leaders in one to two priority Asian markets to understand strategic priorities, business challenges, and opportunities where AI can create measurable value.
  • Assess regional needs against the existing global AI capability portfolio, identifying where current platforms, models, products, and reusable components can be deployed or adapted.
  • Identify capability gaps that require new AI solutions and help shape the roadmap, business case, requirements, and execution approach for those investments.
  • Translate opportunities in the initial KFC markets into a focused portfolio of executable AI initiatives, balancing business value, scalability, technical feasibility, speed to impact, and reuse across markets.
  • Mobilize and coordinate Data Science and Machine Learning Engineering resources across the United States, Colombia, and India to execute priority AI initiatives in the selected markets, ensuring clear ownership, delivery plans, and stakeholder alignment.
  • Work closely with the Global AI Leader and regional AI leaders to align market execution with global AI strategy, architecture, standards, governance, and reusable capabilities.
  • Serve as a bridge between market and regional stakeholders and global AI teams, ensuring local requirements are understood, solutions are appropriately localized, and delivery progress and risks are communicated clearly.
  • As the Asia AI portfolio grows, progressively expand execution responsibility to additional markets and help establish and manage local Data Science and Machine Learning Engineering capability in the region, likely based in Vietnam. The scope of this mandate is expected to grow with demonstrated business value and delivery maturity.

Monitoring & Automation:

  • Design, develop, and maintain monitoring systems to proactively detect anomalies, performance degradation, and system failures.
  • Implement alerting and reporting frameworks to ensure timely intervention and communication of system health metrics.
  • Advocate for automation in operational workflows to minimize manual intervention and improve efficiency.

Metrics & Reporting:

  • Define key operational metrics (e.g., mean time to resolution, issue volume trends, uptime) and deliver regular reports to stakeholders.
  • Continuously evaluate the performance of AI operations and recommend process improvements to enhance system reliability and support quality.

Qualifications

  • Bachelor's or Master's degree in Computer Science, AI, Engineering, or a related field.
  • 5+ years of experience in AI/ML operations, MLOps, or DevOps, with at least 2 years in a leadership role.
  • Proven experience managing 24/7 operations teams and establishing operational processes at scale.
  • Strong understanding of AI/ML systems, pipelines, and monitoring frameworks.
  • Proficiency in tools and platforms for monitoring and incident management (e.g., Prometheus, Grafana, Datadog, PagerDuty).
  • Excellent problem-solving and troubleshooting skills with a hands-on approach to resolving critical issues.
  • Strong English language proficiency, with excellent written and verbal communication skills and the ability to effectively collaborate with technical and non-technical stakeholders across global teams.
  • Demonstrated ability to partner with senior business and technology leaders to identify, assess, prioritize, and shape AI or data-driven opportunities.
  • Experience leading execution across distributed Data Science, Machine Learning Engineering, AI Engineering, or related technical teams in multiple countries or regions.
  • Strong program leadership skills with the ability to translate ambiguous business needs into clear use cases, requirements, delivery plans, and measurable outcomes.
  • Experience in developing and tracking SLAs, operational metrics, and playbooks.
  • Familiarity with cloud platforms (AWS, GCP, or Azure) and container orchestration tools (Kubernetes, Docker).

Preferred Qualifications:

  • Experience with automation tools and scripting languages (e.g., Python, Bash) for operational workflows.
  • Knowledge of machine learning lifecycle management tools (e.g., MLflow, SageMaker, Kubeflow).
  • Previous experience in incident response for AI/ML or data-intensive systems.
  • Experience supporting onboarding processes for large-scale systems or distributed locations (e.g., stores or franchises).

What We Offer:

  • 100% salary during the probation period
  • 18 days of annual leave per year
  • Five “Recharge Days” – extra days in addition to company holidays
  • One day of paid leave for your birthday
  • 2 days WFH/ week
  • Half-day Fridays
  • Full salary insurance
  • 13th-month bonus
  • Advanced health insurance (Generali)
  • Regular team and engagement activities
  • LinkedIn learning and training courses, based on company policy
  • MacBook and monitors

Similar roles