Senior Site Reliability Engineer
EROAD Auckland, Auckland, New Zealand
IT Services and IT Consulting · 201-500 employees
Applying here? Try the free cover letter tool — paste this posting and your résumé, no account needed.
About the role
The Senior Site Reliability Engineer will design, operate, and improve large-scale distributed SaaS platforms while driving root-cause analysis and automation. They will also mentor teammates and optimize infrastructure costs within an autonomous Agile team.
What they look for
Requirements
Candidates must have deep hands-on experience operating Kubernetes and public cloud platforms like Azure or AWS in production. Proficiency in Infrastructure-as-Code, CI/CD tooling, and scripting languages such as Ruby, Bash, or Python is required.
Benefits
Full description
ABOUT THE COMPANY
EROAD is a global fleet management technology company and a genuine Kiwi tech success story, listed on both the NZX and ASX and growing across New Zealand, Australia, the Philippines and the USA. We're hiring a Senior Site Reliability Engineer to help maintain and evolve the SaaS platforms that our fleet operator customers rely on every day.
WHAT YOU'LL DO
This is a hands-on senior role with real influence over how our platforms run and improve.
- Contribute to solution design for large-scale, customer-facing distributed systems
- Operate and support production environments as part of an on-call roster
- Drive root-cause analysis when things go wrong and turn findings into lasting fixes
- Build automation that reduces manual toil and improves reliability
- Improve monitoring and observability across the platform
- Identify and deliver cost optimisation opportunities across the infrastructure
- Mentor teammates and help lift the team's overall SRE practice
- Work within an autonomous, self-managed Agile platform team aligned to one of EROAD's SaaS ecosystems, reporting to the Domain Chapter Lead
WHAT YOU NEED
To thrive in this role you'll bring solid, hands-on production experience across the following.
- Deep hands-on experience operating Kubernetes (AKS or EKS) in production at scale
- Hands-on experience with a public cloud platform, Azure or AWS
- Alignment to one cloud ecosystem, Azure/Windows or AWS/Linux, is fine, you don't need both
- Experience with Terraform or similar Infrastructure-as-Code tooling
- Experience with CI/CD tooling such as GitHub Actions, Azure DevOps, Concourse or similar
- Practical scripting ability in Ruby, Bash, Python or similar
- Experience operating and managing complex, customer-facing, multi-tier distributed production systems
- Exposure to monitoring, alerting or visualisation tools such as Grafana, Sumo Logic or Datadog
WHY JOIN
This is a chance to take real ownership of reliability and performance for platforms operating at national and global scale.
- Senior, hands-on individual contributor role with genuine scope to shape solution design and drive improvement, not just keep the lights on
- Mentoring responsibilities that build your leadership profile alongside your technical depth
- Be part of a high-growth, innovative technology company making a real difference for customers across multiple countries
- Work alongside a talented and collaborative team in an autonomous, self-managed platform team
- Continuous learning support, including EAP offerings and AI tooling to help you grow
- Competitive salary and benefits package
- A multicultural organisation that genuinely values diversity
ABOUT THE COMPANY
EROAD is a global fleet management technology company and a genuine Kiwi tech success story, listed on both the NZX and ASX and growing across New Zealand, Australia, the Philippines and the USA. We're hiring a Senior Site Reliability Engineer to help maintain and evolve the SaaS platforms that our fleet operator customers rely on every day.
WHAT YOU'LL DO
This is a hands-on senior role with real influence over how our platforms run and improve.
- Contribute to solution design for large-scale, customer-facing distributed systems
- Operate and support production environments as part of an on-call roster
- Drive root-cause analysis when things go wrong and turn findings into lasting fixes
- Build automation that reduces manual toil and improves reliability
- Improve monitoring and observability across the platform
- Identify and deliver cost optimisation opportunities across the infrastructure
- Mentor teammates and help lift the team's overall SRE practice
- Work within an autonomous, self-managed Agile platform team aligned to one of EROAD's SaaS ecosystems, reporting to the Domain Chapter Lead
WHAT YOU NEED
To thrive in this role you'll bring solid, hands-on production experience across the following.
- Deep hands-on experience operating Kubernetes (AKS or EKS) in production at scale
- Hands-on experience with a public cloud platform, Azure or AWS
- Alignment to one cloud ecosystem, Azure/Windows or AWS/Linux, is fine, you don't need both
- Experience with Terraform or similar Infrastructure-as-Code tooling
- Experience with CI/CD tooling such as GitHub Actions, Azure DevOps, Concourse or similar
- Practical scripting ability in Ruby, Bash, Python or similar
- Experience operating and managing complex, customer-facing, multi-tier distributed production systems
- Exposure to monitoring, alerting or visualisation tools such as Grafana, Sumo Logic or Datadog
WHY JOIN
This is a chance to take real ownership of reliability and performance for platforms operating at national and global scale.
- Senior, hands-on individual contributor role with genuine scope to shape solution design and drive improvement
- Mentoring responsibilities that build your leadership profile alongside your technical depth
- Be part of a high-growth, innovative technology company making a real difference for customers across multiple countries
- Work alongside a talented and collaborative team in an autonomous, self-managed platform team
- Continuous learning support, including EAP offerings and AI tooling to help you grow
- Competitive salary and benefits package
- A multicultural organisation that genuinely values diversity
Similar roles
-
Site Reliability Engineer I
American Express Bengaluru, Karnataka, India
-
Site Reliability Engineering Lead
trivago Düsseldorf, North Rhine-Westphalia, Germany
-
Lead SRE
Apple Shanghai, Shanghai, China
-
Site Reliability Engineer
Ensono pune, Maharashtra, India
-
Specialist, Application Support - GCP Site Reliability Engineering
NielsenIQ Chennai, Tamil Nadu, India
-
SRE Manager (Traffic and Secure Networking)
Apple Shanghai, Shanghai, China