Site Reliability Engineer II
Yum! · Ho Chi Minh City, Vietnam
Restaurants · 10,001+ employees
About the role
The role involves providing reliability coverage and facilitating handoffs between global shifts for restaurant technology. It focuses on transforming reliability practices from reactive incident response to a platform engineering model using automation and self-service tools.
What they look for
Requirements
Candidates must have 5+ years of experience in site reliability engineering or production operations with proficiency in cloud providers and observability tooling. Strong communication skills and technical depth in automation, Kubernetes, and infrastructure as code are required.
Full description
- 5+ years in site reliability engineering, DevOps, infrastructure, or production operations roles
- Hands-on credibility across incident response, observability, and automation, with the technical depth to guide Level 6-7 engineers
- Experience operating in shift-based, on-call, or follow-the-sun coverage models
- Working knowledge of at least one major cloud provider (AWS preferred) and modern observability tooling (e.g., Datadog, Prometheus, Grafana)
- Proficiency in at least one scripting or programming language sufficient to review and guide automation work
- Understanding of SLI/SLO frameworks and reliability engineering fundamentals
- Strong written and verbal English communication skills for cross-region collaboration with US and India teams