ZURU

Site Reliability Engineer

ZURU Auckland, Auckland, New Zealand

Manufacturing · 1,001-5,000 employees

13 h ago
sre Senior (5-10 yrs) Full-time New Zealand
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

You will design and build the reliability function for ZURU's internal application estate, including establishing automated graduation gates and runtime foundations. You will also implement observability, define SLOs, and scale incident response processes to ensure system stability.

What they look for

AWS Observability Automation Terraform Lambda ECS EKS Temporal Open Telemetry SLO definition Incident response Pipeline management Reliability engineering System architecture

Requirements

The ideal candidate is an experienced engineer with strong AWS skills and a proven ability to automate toil rather than relying on manual runbooks. You must have hands-on experience with observability tools and a track record of defining SLOs that drive technical decision-making.

Benefits

Health benefits Competitive remuneration

Full description

About ZURU

ZURU is on a mission to disrupt across industries, challenge the status quo and catalyst change through radical innovation and automation advances. This is in play in different pillars of the company: ZURU Toys are re-imagining what it means to play; ZURU Tech is shaping a better future by leading the next building revolution; ZURU Edge is pioneering new generation FMCG brands to better serve modern consumers.

 

Founded in 2003 by EY Entrepreneur of the Year and World Entrepreneur Hall of Fame brothers Nick and Mat Mowbray, ZURU has quickly grown to a team of over 5000 direct and indirect members across more than 30 international locations.

 

One of the largest toy companies in the world, globally recognised and award-winning brands include Bunch O Balloons, Mini Brands, XSHOT, Rainbocorns and Smashers. Our global FMCG brands include MONDAY Haircare, Rascals, NOOD, BONKERS, Gumi Yum Surprise, and more!

For more information, visit www.zuru.com.

Position Overview

This is the founding SRE role for ZURU's internal application estate — the reliability function that lets an application outlive the engineer who built it. ZURU's internal engineers sit with the business units they serve and build fast. This role is how what they build stays up. Working as one of a pair — Auckland and China, covering the estate follow-the-sun — you're not joining a reliability function. You're designing it.

Position Impact

At 6 months, the graduation gate runs as a pipeline job, the runtime foundations are hardened and published so every new build inherits them, and the first applications have moved off the engineers who built them onto a supported rota.

At 12 months and beyond, the estate gets quieter rather than busier. Applications graduate and retire each quarter, ungraduated applications per engineer stays inside the ceiling the whole model depends on, and mean time to restore and change-fail rate both trend down — because repeat failures are fixed in the shared capability package rather than in a runbook.

Who you are:

An engineer who has carried a pager for systems other people wrote — and made that sustainable rather than heroic. You're strong on AWS, hands-on with observability, and you automate your way out of toil rather than adding a runbook entry. You've written an SLO that actually changed a decision. You don't need a reliability function to already exist — you're here to build one.

Key Responsibilities

The Graduation Gate

  • Own the automated checks an application passes to join the supported estate: on the paved path, observability wired, an SLO defined, alerts routed to a rota, reproducible from a clean clone.
  • Run them as a pipeline job, not a meeting.

Runtime Foundations

  • Harden and publish AWS across Lambda, ECS, EKS, and Terraform — with Temporal for durable workflows — so every new build inherits them from day one.

Observability by Default

  • Wire Open Telemetry from day one into a single telemetry layer.
  • Define SLOs with each application's business owner in terms they would notice breaching.

Agent-First Incident Response

  • Design the loop where an agent drafts the fix and a human approves it.
  • Build incident response that scales without scaling the team.

Fixes That Stay Fixed

  • Push every repeat failure into the capability package every repository consumes — so every future build inherits the fix, rather than adding a runbook entry.
  • Make the estate quieter over time, not louder.

LIFE@ZURU

ZURU is on a quest to reimagine tomorrow. Founded in 2003, ZURU Group has rapidly grown and now spans three core divisions—ZURU Toys, ZURU Edge (consumer goods) and ZURU Tech (construction).

At ZURU, we have cultivated a high-performing culture that encourages excellence. Our team works towards ambitious goals, learning, performing, and improving together, all while having fun. We empower talented individuals to do their best work every day.

At ZURU, you get out what you put in. You are responsible for driving your own career and we provide the platform to achieve it. As ZURU is on such a fast growth trajectory, there are opportunities here that you won’t find anywhere else.

Get to know us a little better by checking out @lifeatzuru on Instagram or www.zuru.com.

WHAT WE OFFER

  • 🌱 Culture for Growth
  • 🧘 Health Benefits
  • 🌎 Global Opportunities
  • 💡 Surrounded by an A Player Team
  • 💰 Competitive Remuneration
  • 🍓 Lots of fresh fruit, coffee, pals fridge and more (optional)

ZURU – Reimagining tomorrow 🚀

Similar roles