Senior Site Reliability Engineer
AoFrio Auckland, Auckland, New Zealand
Software Development · 51-200 employees
About the role
You will lead the design, delivery, and reliability of AWS infrastructure, including container platforms and production databases. You will also define monitoring standards, manage CI/CD strategies, and mentor junior engineers to ensure high system availability.
What they look for
Requirements
The role requires extensive experience in SRE or DevOps, with deep expertise in AWS, container orchestration (ECS/EKS), and infrastructure-as-code. Strong fundamentals in Linux, networking, and observability tools are essential for managing production-scale distributed systems.
Benefits
Full description
Welcome to our World of Cold! At AoFrio, we are global leaders in providing IoT solutions to the food and beverage industry. Our innovative technology and dedicated team have positioned us at the forefront of our market.
Every AoFrio controller, monitor and gateway in the field depends on a platform underneath it: the AWS infrastructure, the databases, the pipelines that ship code safely at 2am if it has to. You'll own that platform end to end from how it's built, how it scales, how fast we know when something's wrong to how quickly we fix it. You'll set the standard other engineers follow and you'll work through complex, ambiguous problems with minimal oversight.
You will do:
- Lead the design and delivery of AoFrio's AWS infrastructure — containers on ECS and EKS, microservices and distributed systems — built for reliability and scale
- Own major reliability and infrastructure workstreams end to end, from architecture through to operation
- Define how we monitor, alert and get visibility into our platforms, and set the standards for automation and infrastructure-as-code
- Set our CI/CD, GitOps and release strategy — including blue/green and canary deployments — and drive the team to adopt it
- Own the reliability of our production databases: high availability, backup and recovery, and performance
- Lead the response to major incidents, own root cause analysis, and drive fixes through to done; take your turn in the on-call rotation
- Be the escalation point when a platform reliability problem has no playbook yet
- Mentor and review the work of intermediate and junior engineers, and set the technical bar others measure against
You'll need:
- Proven expertise in working as a Site Reliability Engineer, DevOps Engineer or similar, including recent hands-on experience operating production cloud infrastructure at scale
- Deep AWS expertise, including architecture and cost optimisation, backed by strong Linux and networking fundamentals (VPC, DNS, load balancing, security groups)
- Extensive production experience designing and operating container platforms on ECS and EKS, including microservices and distributed systems
- Strong infrastructure-as-code and CI/CD experience (AWS CDK, CloudFormation, GitOps with ArgoCD), including setting standards and release strategies such as blue/green and canary
- Experience designing high-availability, backup and recovery, and performance strategies for self-hosted production databases
- Proven experience designing monitoring, alerting and observability using tools such as Amazon CloudWatch, Prometheus and Grafana
- Proven ability to design scalable, reliable and secure systems from first principles, with the communication skills to explain the reasoning to non-technical stakeholders
- Beneficial: experience leading the use of AI tools and agents in SRE practice, such as incident diagnosis and runbook automation
- Beneficial: tertiary qualification in computer science, software engineering or a related field, or equivalent practical experience
Who we are We're leaders in hardware-enabled SaaS for commercial refrigeration. Our controllers, monitors and gateways sit inside coolers for Coca-Cola, PepsiCo and AB InBev, and the software on top turns them into data our customers actually act on. Energy saved, stock protected, technicians sent to the right site.
We run supply chain, logistics and R&D out of our Auckland HQ, with offices in four countries and people in twelve. It's a small enough team that your work is visible, and a global enough business that it matters.
What it's like to work here We're a global business with the corporate parts that come with that and we're also a place where plenty of things don't have a process yet. Some areas are polished, others you'll end up shaping yourself. If you like a bit of room to do that, you'll be fine here.
The Auckland office is properly international and the personalities are all over the place. Quiet, loud, neurodiverse, people at every stage of life. Nobody has to fit a particular mould to do well here. On Thursdays you might also run into a dog.
On AI: we use it, we pay for premium tools and we'd rather you spent your time on hard problems than on boilerplate. If you want to bring AI into how you work or into what we build, you won't need to make a business case for it first.
What you'll enjoy at AoFrio:
- Southern Cross WellbeingOne health insurance, fully covered by AoFrio.
- 16 weeks of AoFrio top-up payments for primary carers, on top of government paid parental leave.
- 4 weeks of fully paid leave for secondary caregivers. Statutory partner's leave in New Zealand is unpaid. Ours isn't.
- Free annual flu vaccinations.
- Professional development funds to spend on development & growth.
- Genuine flexible working, including work from home when needed.
- Work from anywhere in the world for up to 4 weeks a year, after 12 months with us (depending on the role).
- A paid volunteer day each year, for a cause you care about.
- Free parking and free EV charging.
- Fruit, cereal, snacks, barista coffee machine and cold drinks.
- Birthday celebrations and social events worth actually turning up to.
- AoWLead, our internal network supporting women at AoFrio.
Diversity at AoFrio We're building a team that reflects the world our products end up in. We actively welcome women in tech, and anyone who doesn't see themselves in the usual tech job ad. Whoever you are, you'll be supported here.
Tell us what you need to do your best work, whether that's an adjustment to the interview process, a different working pattern or something we haven't thought of. Ask us! We'll work it out.
Similar roles
-
Site Reliability Engineer
H&M Group Bengaluru, Karnataka, India
-
Senior Infrastructure SRE
Apple Cupertino, California, United States
-
Senior Site Reliability Engineer Lead
Akamai Cambridge, Massachusetts, United States · $121K–$219K/yr
-
Site Reliability Engineer
SGX Singapore, Singapore
-
Director, Site Reliability Engineering
Jobgether Canada · $244K/yr
-
Site Reliability Specialist - 12 months
Co-operators Guelph, Ontario, Canada · CA$61K–CA$101K/yr