Site Reliability Engineer II
Yum! Irvine, California, United States · $105K–$132K/yr
Restaurants · 10,001+ employees
About the role
The Site Reliability Engineer will maintain digital system availability and performance while automating processes to reduce toil. They will collaborate with cross-functional teams to troubleshoot incidents, conduct root cause analysis, and improve system observability.
What they look for
Requirements
Candidates must have at least 2 years of experience in SRE with a strong focus on observability and automation. A bachelor's degree in a technical field or equivalent experience is required, along with proficiency in cloud-native environments and scripting languages.
Benefits
Full description
Who is Taco Bell?
Taco Bell was born and raised in California and has been around since 1962. We went from selling everyone’s favorite Crunchy Tacos on the West Coast to a global brand with 7,500+ restaurants, 350 franchise organizations, that serve 42+ million fans each week around the globe. We’re not only the largest Mexican-inspired quick service brand (QSR) in the world, we’re also part of the biggest restaurant group in the world: Yum! Brands.
Much of our fan love and authentic connection with our communities are rooted in being rebels with a cause. From ensuring we use high-quality, sustainable ingredients to elevating restaurant technology in ways that haven’t been done before… we will continue to be inclusive, bold, challenge the status quo and push industry boundaries.
We’re a company that celebrates and advocates for different, has bold self-expression, strives for a better future, and brings the fun while we’re at it. We fuel our culture with real people who bring unique experiences. We inspire and enable our teams and the world to Live Más.
At Taco Bell, we’re Cultural Rebels. Want to join in on the passion-fueled fun? Learn more about the career below.
About The Team
The Digital Site Reliability Engineering (SRE) team supports one of the fastest-growing lines of business at one of the hottest Quick Service Restaurants as part of the largest group of restaurants in the world.
We support both internal and external customers by ensuring that our digital systems, including mobile and web ordering, operate as smoothly and seamlessly as possible.
As part of the SRE team, Site Reliability Engineers work closely with Software Developers, Platform Engineering teams, and other stakeholders to help ensure our digital systems are available and performing as expected.
Our goal is to deploy and maintain observability, reduce toil through automation, identify and remediate defects, and continuously improve both the environment and our own capabilities. We also ensure we have, or build, the tools necessary to identify root causes and answer questions we had not yet anticipated.
Many team members are remote/distributed across the US, while others reside near our Irvine campus Headquarters, and come into the office on a hybrid schedule.
About The Role
We’re looking for someone who can own a problem from start to finish — someone who is comfortable digging into the code, identifying the issue, and proposing a fix. This person will become a subject matter expert within the Taco Bell Digital space.
Fundamentally, we’re looking for someone who lives and breathes observability and automation, and who is passionate about continuously raising the bar.
We want someone who listens well, communicates clearly and effectively, drives alignment, and helps engineering teams improve their processes.
As a fierce advocate for the customer, your job would take you into troubleshooting issues and incidents, building out new dashboards and alerts, finding out the answers to help us get to the root cause of problems, and ultimately fixing them for good.
Responsibilities
The Day-to-Day:
- Create, update, or automate internal business processes or tools to reduce toil and improve team productivity.
- Understand and monitor the Taco Bell Digital ecosystem for performance, availability, and accuracy of transactional data.
- Communicate and collaborate with both technical and non-technical stakeholders on issues, upcoming changes, and updates to system health.
- Perform final validation tests on various mobile and web-based applications, reporting on, and offering feedback on areas for improvement.
- Build expertise in serverless infrastructure and initiatives while also learning aspects of modern SRE practices and terms, such as SLIs, SLOs, Observability, toil, and incident response with blameless postmortems.
- Work with and adopt Agile practices while participating in a 24/7 on-call rotation.
- Collaborate within the team and with cross-functional partners on high-impact business issues that affect revenue and brand reputation.
Qualifications
Is this you?
Requirements:
- Bachelor’s degree in computer science, engineering, OR a related field, OR equivalent work experience.
- At least 2+ years of experience in the SRE space, with a focus on observability and automation
- Familiarity with SRE core principles (e.g. SLO, SLA, SLI, Error Budget, etc.)
- Hands-on experience creating monitors, dashboards, SLOs, and other observability capabilities
- Experience with logging solutions or platforms such as DataDog, CloudWatch log insights, etc.
- Understanding of incident management practices, including leading bridge calls, conducting RCAs, and facilitating postmortems
- Familiarity with modern observability practices and tools, such as distributed tracing, APM, OpenTelemetry, and RUM
- Excellent communication and collaboration skills, with the ability to work effectively in a fast-paced environment as a member of a team
- A fundamentally complete understanding of Observability principles (not just monitoring) + experience using tools like DataDog, Lumigo, CloudWatch, Incident.io, PagerDuty or similar
- General level understanding of Agile methods such as Kanban, Scrum, etc.
- Advanced troubleshooting skills
- A curious mindset and the desire to always keep learning
- Proactive self-starter capable of operating autonomously
- Ability to participate in an on-call rotation
- Proficiency with SQL (intermediate)
- Intermediate understanding of JavaScript, Python, Go, or TypeScript
- Fundamental knowledge of AWS services commonly used in serverless and cloud-native environments, including Lambda, API Gateway, Fargate, S3, DynamoDB, and EventBridge
Preferred:
- Skills surrounding software development (Git, CI/CD, reading and writing code) - preferably in JavaScript/TypeScript and/or Python and with tools like VS Code, Gitlab CI/CD or GitHub Actions
- Experience with building automations for internal business processes, including AI-embedded workflows
- Comfort and familiarity with common Unix-like shells (bash, zsh, etc)
- Experience with Retool
- Experience with data and observability platforms: FullStory, Amplitude, Embrace
- Familiarity with AI agentic tools like Claude Code, Codex, Cursor, etc.
- Prior use of issue tracking systems such as Jira
- High proficiency with AWS serverless services, including Lambda, API Gateway, Fargate, S3, DynamoDB, and EventBridge.
- Experience with Infrastructure as Code (IaC) tools such as Terraform, Pulumi, or CloudFormation
- Knowledge of Akamai tools, processes, and SOCC engagement
- Proficiency with deploying AI via AWS Bedrock or Agentcore
Work-Hard, Play-Hard:
- Hybrid work schedule and year-round flex day Friday (half day)
- Onsite childcare through Bright Horizons
- Onsite dining center and game room (yes, there is a Taco Bell inside the building)
- Onsite dry cleaning, laundry services, carwash,
- Onsite gym with fitness classes and personal trainer sessions
- Up to 4 weeks of vacation per year plus holidays and time off for volunteering
- Tuition reimbursement and education benefits
- Generous parental leave for all new parents and adoption assistance program
- 401(k) with a 6% matching contribution from Yum! Brands with immediate vesting
- Comprehensive medical & dental including prescription drug benefits and 100% preventive care
- Discounts, free food, swag and… honestly, too many good benefits to name
Salary Range $105,000 to $132,200 annually + bonus eligibility + equity (if applicable) + benefits
The above represents the expected salary range for this job requisition. Ultimately, in determining your pay, we'll consider your location, experience, and other job-related factors.
At Taco Bell, we Live Más and invite you to do the same. Take a seat at our table. Bring your voice. Bring you, just as you are, a Cultural Rebel. We want you to be your best self!
Taco Bell is proud to be an equal opportunity employer and is committed to equity, inclusion, and belonging for all dimensions of diversity. We do not discriminate based on race, color, religion, sex, sexual orientation, gender identity, national origin, veteran status, disability status, age, or any other protected characteristic.
Taco Bell is committed to working with and providing reasonable accommodation to applicants with disabilities or special needs.
US Job Seekers/Employees - Click here to view the “Know Your Rights” poster and supplement and the Pay Transparency Policy Statement. Employment eligibility to work with Taco Bell in the U.S. is required as the company will not pursue visa sponsorship for this position.
Similar roles
-
Site Reliability Engineer II (SRE)
Xometry Quinte West, Ontario, Canada · $135K–$155K/yr
-
TSP SRE expert
Berge Group Gothenburg, Nebraska, United States
-
Site Reliability Engineer
ASM Research United States
-
Cloud Systems Engineer - Site Reliability
TherapyNotes.com Philadelphia, Pennsylvania, United States · $110K–$150K/yr
-
Site Reliability Manager, Data Center Networking, SRE
Google Waterloo, Ontario, Canada · CA$216K–CA$221K/yr
-
Site Reliability Engineer II
Akamai Bengaluru, Karnataka, India