Assoc. Manager, Site Reliability Engineering
KFC Europe Ho Chi Minh City, Vietnam
Restaurants · 10,001+ employees
About the role
The Associate Manager leads the Vietnam-based SRE team, overseeing daily operational discipline, incident response, and shift handoffs. They are responsible for mentoring early-to-mid career engineers and fostering a culture of ownership within a global follow-the-sun coverage model.
What they look for
Requirements
Candidates must have at least 5 years of experience in SRE or DevOps and 1 year of people management experience. Proficiency in cloud infrastructure, observability tools, and scripting languages is required, along with strong English communication skills.
Full description
- 5+ years of experience in site reliability engineering, DevOps, infrastructure, or production operations roles
- 1+ years of people management experience, or 2+ years as a senior technical lead with demonstrated coaching and delivery ownership
- Hands-on credibility across incident response, observability, and automation, with the technical depth to guide Level 6-7 engineers
- Experience operating in shift-based, on-call, or follow-the-sun coverage models
- Working knowledge of at least one major cloud provider (AWS preferred) and modern observability tooling (e.g., Datadog, Prometheus, Grafana)
- Proficiency in at least one scripting or programming language sufficient to review and guide automation work
- Understanding of SLI/SLO frameworks and reliability engineering fundamentals
- Strong written and verbal English communication skills for cross-region collaboration with US and India teams
Responsibilities
- 5+ years of experience in site reliability engineering, DevOps, infrastructure, or production operations roles
- 1+ years of people management experience, or 2+ years as a senior technical lead with demonstrated coaching and delivery ownership
- Hands-on credibility across incident response, observability, and automation, with the technical depth to guide Level 6-7 engineers
- Experience operating in shift-based, on-call, or follow-the-sun coverage models
- Working knowledge of at least one major cloud provider (AWS preferred) and modern observability tooling (e.g., Datadog, Prometheus, Grafana)
- Proficiency in at least one scripting or programming language sufficient to review and guide automation work
- Understanding of SLI/SLO frameworks and reliability engineering fundamentals
- Strong written and verbal English communication skills for cross-region collaboration with US and India teams
Qualifications
- Experience building or standing up a new team, site, or shift operation
- Experience managing engineers across the early-to-mid career range with a track record of promotions or level progression
- Kubernetes, container orchestration, and infrastructure as code experience (e.g., Terraform)
- Familiarity with AI-assisted operations tooling and automation-first reliability approaches, including auto-healing and auto-remediation patterns
- Exposure to platform engineering and internal developer platform concepts: self-service tooling, developer portals (e.g., Port, Backstage), GitOps
- Experience in multi-region or globally distributed team models
- Relevant certifications (AWS, CKA, or similar)
Similar roles
-
Principal Site Reliability Engineer, Platform Engineering: Dedicated
GitLab Canada · $223K–$380K/yr
-
Senior Site Reliability Engineer (SRE)
UJET United States · $140K–$180K/yr
-
Senior Site Reliability Engineer
Apple Cupertino, California, United States
-
Site Reliability Engineer (High Performance Computing)
SpaceX Hawthorne, California, United States · $125K–$195K/yr
-
Site Reliability Engineer
N26 Barcelona, Catalonia, Spain
-
SRE
Hitachi Solutions pune, Maharashtra, India