Yum!

Site Reliability Engineer II

Yum! Ho Chi Minh City, Vietnam

Restaurants · 10,001+ employees

4 d ago
sre Senior (5-10 yrs) Full-time United States
Log in to apply, save this posting, or score it against your profile with AI.

About the role

The Site Reliability Engineer II will manage incident response, observability, and automation while providing technical guidance to engineering teams. They will operate within shift-based or follow-the-sun models to ensure system reliability and support cross-region collaboration.

What they look for

Site Reliability Engineering DevOps Infrastructure Incident Response Observability Automation AWS Datadog Prometheus Grafana Kubernetes Terraform GitOps Platform Engineering SLI/SLO Frameworks Scripting

Requirements

Candidates must have at least 5 years of experience in SRE, DevOps, or production operations and proficiency in cloud platforms and observability tools. Strong scripting skills and an understanding of reliability engineering fundamentals are required to support global infrastructure.

Full description

  • 2+ years of experience in SRE, DevOps, production support, or infrastructure engineering roles
  • Hands-on experience with monitoring and observability tooling (e.g., Datadog, Prometheus, Grafana, CloudWatch, or similar)
  • Working knowledge of at least one major cloud provider (AWS preferred)
  • Proficiency in at least one scripting or programming language (e.g., Python, Bash, Go) for automation, with demonstrated examples of automating away manual operational work
  • Experience participating in incident response and on-call or shift-based operations
  • Understanding of SLI/SLO concepts and reliability engineering fundamentals
  • Ability to work follow-the-sun shift rotations, including structured handoffs with teams in other regions
  • Strong written and verbal English communication skills for cross-region collaboration

Similar roles