Senior Site Reliability Engineer - Observability
Westpac Group Sydney, New South Wales, Australia
Financial Services · 10,001+ employees
About the role
The role involves embedding observability into architecture, engineering design, and production operations for critical banking platforms. It also requires participation in a 24x7 on-call roster for incident response and service restoration.
What they look for
Requirements
Candidates must have advanced hands-on experience with Elastic Stack, AWS observability tools, and Splunk. Strong knowledge of SRE principles, distributed tracing, and major incident management is essential.
Benefits
Full description
Create your best future and join Westpac as a Senior Site Reliability Engineer.
What’s the role?
We are seeking an experienced Senior Site Reliability Engineer with deep expertise in observability engineering, cloud monitoring, reliability engineering and AI-enabled operations to support critical Payments, Treasury and Digital Banking platforms.
The platform is at a critical stage of growth, with additional payment types being onboarded and liquidity capabilities expanding into the core platform. This role will ensure observability is embedded into architecture, engineering design and production operation, including logging, metrics, distributed tracing, application performance monitoring, alerting, service health dashboards and end-to-end visibility of customer and payment journeys.
The role includes participation in a 24x7 production support and on-call roster for critical payment systems, with after-hours incident response, service restoration and escalation support when required and in accordance with organisational roster arrangements.
What do I need?
- Advanced hands-on Elastic Stack experience, particularly Elasticsearch and Kibana.
- Strong Amazon CloudWatch and AWS observability experience.
- Strong Splunk Enterprise and/or Splunk Observability experience.
- OpenTelemetry, distributed tracing, APM, telemetry pipelines, dashboards and alert engineering.
- SRE principles, SLIs, SLOs, error budgets, reliability and performance engineering.
- Major incident management, problem management and root cause analysis.
- Experience supporting business-critical production services and improving MTTD and MTTR.
- Ability and willingness to participate in a 24x7 on-call roster for critical payment systems.
Why join us?
We’re obsessed with becoming our customers' #1 banking partner for life and we’re looking for people who are passionate about helping us achieve that goal. In return, we’re committed to making Westpac the best place to work in the country. Here are just a few of the ways we’re already doing that:
- Special offers on banking products and discounts from top brands, including generous employee-only mortgage rates!
- Flexible work arrangements to help you achieve a greater work/life balance, and a variety of leave options including Culture, Lifestyle and Wellbeing leave.
- Tailored learning and development opportunities to help your grow your career within the bank.
- Lots of opportunities to ‘give back’ to the Community by getting involved in our many volunteering initiatives.
Create your future today
To get started, simply click on the APPLY or APPLY NOW button
We’re all about creating a supportive and inclusive community. We welcome everyone – no matter your age, gender, background, or abilities. We also provide additional support to welcome our veterans, Indigenous Australians and neurodiverse community.
If you need any adjustments during the recruitment process, you can find more information and contact details on our FAQs and how to contact us page, under the ‘Diversity, sustainability and flexibility’ section.
We may close this job advertisement earlier than the advertised closing date if suitable candidates are identified. We encourage you to apply as soon as possible.
#LI - Hybrid.
Similar roles
-
Sr. Site Reliability Engineer, AI Infrastructure (Starshield)
SpaceX Washington, District of Columbia, United States · $165K–$265K/yr
-
Sr. Site Reliability Engineer, AI Infrastructure (Starshield)
SpaceX Redmond, Washington, United States · $165K–$270K/yr
-
Site Reliability Engineer, AI Infrastructure (Starshield)
SpaceX Redmond, Washington, United States · $125K–$200K/yr
-
Senior Site Reliability Engineer
Jobgether India
-
Analyst, Engineer (SRE) - Home Ownership - Broker My Home
NAB Hanoi, Vietnam
-
SRE Sênior
Jobgether Brazil