Westpac Group

Site Reliability Engineer

Westpac Group Sydney, New South Wales, Australia

Financial Services · 10,001+ employees

3 h ago
sre Senior (5-10 yrs) Full-time Australia
Log in to apply, save this posting, or score it against your profile with AI.

About the role

The role involves improving the reliability, resilience, and observability of critical digital banking platforms through automation and modern engineering practices. You will lead the adoption of AI-driven operations and manage incident response to ensure high availability and performance of customer-facing services.

What they look for

Site Reliability Engineering AWS Azure Kubernetes Infrastructure as Code CI/CD Python Java Go Observability Splunk Dynatrace Prometheus AI-driven operations Incident management Cloud-native technologies

Requirements

Candidates must have proven experience in Site Reliability Engineering or DevOps with strong proficiency in cloud-native technologies, Kubernetes, and automation scripting. Experience with AI-assisted engineering tools and a background in financial services or mission-critical systems is highly desirable.

Benefits

Banking product discounts Generous employee-only mortgage rates Flexible work arrangements Culture leave Lifestyle leave Wellbeing leave Tailored learning and development opportunities Volunteering initiatives

Full description

What’s the role?

Westpac One Digital is seeking a highly motivated Site Reliability Engineer to help build, operate and continuously improve the reliability, resilience, observability and performance of critical digital banking platforms. This role combines strong Site Reliability Engineering practices with modern AI-driven engineering to shape Agentic SRE and intelligent operations across Westpac One Digital, ensuring highly available, scalable and secure customer-facing services while driving automation, operational excellence and continuous improvement.

Key Responsibilities

  • Improve service reliability, availability and resilience by defining and managing SLIs, SLOs, Error Budgets, disaster recovery capabilities and operational readiness activities.
  • Design and enhance observability through monitoring, alerting, dashboards, synthetic monitoring, automated health checks and end-to-end visibility across applications, infrastructure and cloud environments.
  • Drive automation and engineering excellence by developing self-healing solutions, operational tooling, platform automation, Infrastructure as Code and reliable CI/CD deployment practices.
  • Participate in major incident management, service restoration and root cause analysis, implementing permanent solutions to reduce recurring issues, customer impact and recovery times.
  • Lead the adoption of AI-driven operations, Agentic SRE capabilities, LLM-powered incident management and GitHub Copilot-enabled engineering practices to improve efficiency and reduce operational toil.

What do I need?

  • Proven experience in Site Reliability Engineering, Production Engineering, DevOps or Platform Engineering, with a strong understanding of distributed systems, cloud-native technologies and modern application architectures.
  • Hands-on experience with AWS and/or Azure, Kubernetes, container platforms, CI/CD pipelines and Infrastructure as Code (IaC) practices.
  • Strong experience designing and supporting observability solutions using tools such as Splunk, Dynatrace, Grafana, Prometheus, OpenTelemetry or equivalent monitoring platforms.
  • Proficiency in software development and automation using languages such as Python, Java, Go, PowerShell or similar scripting and programming technologies.
  • Experience leveraging AI-assisted engineering tools including GitHub Copilot, Microsoft Copilot, Claude, ChatGPT or similar technologies, with an understanding of LLMs, AI agents, prompt engineering, AI governance and responsible AI principles.
  • Demonstrated expertise in operational excellence, including incident management, problem management, change management, root cause analysis, performance engineering, capacity planning, disaster recovery testing and operational readiness.
  • Desirable experience within banking or financial services, including knowledge of APRA CPS 230 Operational Resilience requirements, support of mission-critical digital banking or payments platforms, and exposure to AIOps, AI agents or autonomous operational capabilities.

Why join us?

We’re obsessed with becoming our customers' #1 banking partner for life and we’re looking for people who are passionate about helping us achieve that goal. In return, we’re committed to making Westpac the best place to work in the country. Here are just a few of the ways we’re already doing that:

  • Special offers on banking products and discounts from top brands, including generous employee-only mortgage rates!
  • Flexible work arrangements to help you achieve a greater work/life balance, and a variety of leave options including Culture, Lifestyle and Wellbeing leave.
  • Tailored learning and development opportunities to help your grow your career within the bank.
  • Lots of opportunities to ‘give back’ to the Community by getting involved in our many volunteering initiatives.

Create your future today

To get started, simply click on the APPLY or APPLY NOW button

We’re all about creating a supportive and inclusive community. We welcome everyone – no matter your age, gender, background, or abilities. We also provide additional support to welcome our veterans, Indigenous Australians and neurodiverse community.

If you need any adjustments during the recruitment process, you can find more information and contact details on our FAQs and how to contact us page, under the ‘Diversity, sustainability and flexibility’ section.

We may close this job advertisement earlier than the advertised closing date if suitable candidates are identified. We encourage you to apply as soon as possible.

Similar roles