Yum!

Site Reliability Engineer II

Yum! · Ho Chi Minh City, Vietnam

Restaurants · 10,001+ employees

11 h ago
Senior (5-10 yrs) Full-time Vietnam
Log in to apply, save this posting, or score it against your profile with AI.

About the role

The role involves providing reliability coverage and facilitating handoffs between global shifts for restaurant technology. It focuses on transforming reliability practices from reactive incident response to a platform engineering model using automation and self-service tools.

What they look for

Site Reliability Engineering DevOps Infrastructure Incident Response Observability Automation AWS Datadog Prometheus Grafana Kubernetes Terraform Platform Engineering GitOps Container Orchestration

Requirements

Candidates must have 5+ years of experience in site reliability engineering or production operations with proficiency in cloud providers and observability tooling. Strong communication skills and technical depth in automation, Kubernetes, and infrastructure as code are required.

Full description

  • 5+ years in site reliability engineering, DevOps, infrastructure, or production operations roles
  • Hands-on credibility across incident response, observability, and automation, with the technical depth to guide Level 6-7 engineers
  • Experience operating in shift-based, on-call, or follow-the-sun coverage models
  • Working knowledge of at least one major cloud provider (AWS preferred) and modern observability tooling (e.g., Datadog, Prometheus, Grafana)
  • Proficiency in at least one scripting or programming language sufficient to review and guide automation work
  • Understanding of SLI/SLO frameworks and reliability engineering fundamentals
  • Strong written and verbal English communication skills for cross-region collaboration with US and India teams