HSBC Global Services Limited

Lead Site Reliability Engineer

HSBC Global Services Limited Kowloon, Hong Kong, China

Financial Services · 10,001+ employees

Sep 01 Closes in 7d
sre Principal (10+ yrs) Full-time China
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

Lead the engineering design and delivery of AIOps journeys to improve service resilience and operational effectiveness. Collaborate with cross-functional teams to translate operational pain points into secure, reusable, and supportable engineering capabilities.

What they look for

Site Reliability Engineering AIOps Software Engineering Service Management Python Java Observability Application Performance Management Cloud-native Incident Response Automation API Integration Data Handling Infrastructure Systemic Troubleshooting Technical Leadership

Requirements

Requires 10+ years of experience in software engineering, site reliability engineering, or technology management. Demonstrable expertise in AIOps, observability, and troubleshooting complex incidents in high-pressure production environments is essential.

Benefits

Continuous professional development Flexible working Inclusive and diverse environment

Full description

GCB 4

We are currently seeking a high calibre professional to join our team as a Lead Site Reliability Engineer

In this role you will:

  • Work closely with the Service Management Transformation Lead and wider Service Management community to design and deliver the progressive transformation plan for Service Management practices across the Hong Kong Pillar and Asia & Middle East organisation.
  • Shape engineering-led discussions with Service Management practitioners, SREs, engineers and stakeholders to identify priority AIOps use cases, define solution requirements and develop pragmatic delivery options.
  • Lead the engineering design and delivery of AIOps journeys that improve service resilience and operational effectiveness in day-to-day Service Management work.
  • Co-Build and Co-lead a Team Engineering model in which engineers, Service Management practitioners and SREs co-develop solutions and jointly deliver measurable service outcomes.
  • Translate operational pain points in Incident, Problem, Change and other Service management practices into secure, reusable and supportable engineering capabilities.
  • Lead delivery of AIOps use cases for earlier detection, faster technical diagnosis, better impact assessment, stronger evidence correlation, and more coherent recovery.
  • Drive deep application performance and observability improvements across applications, middleware, databases, infrastructure and platform layers.
  • Establish a clear delivery operating model, including joint backlog prioritisation, decision rights, roadmap management and dependency resolution across teams.
  • Ensure AI-assisted and agentic capabilities are bounded, controlled and auditable, with human validation for consequential production decisions.
  • Partner with platform, cloud, security, data, architecture, Risk and Control teams to align delivery with strategic enterprise capabilities and policy requirements.
  • Contribute technical leadership to the Service Management and SRE practitioner community to improve adoption, consistency and engineering outcomes.
  • Develop and support tooling to facilitate major incident response with evidence-led technical analysis and post-incident deep dives that drive durable remediation and recurrence reduction.
  • Drive a culture of continuous improvement in engineering quality, operational discipline, and service management outcomes.

To be successful you will need:

  • 10+ years of experience in software engineering, site reliability engineering, or technology management roles with strong customer and service focus.
  • Demonstrable hands-on technical expertise in engineering design, delivery and operation of production-critical services.
  • Strong ability to navigate engineering-led discovery and requirements discussions, translating operational needs into clear use cases, technical requirements, controls and measurable outcomes.
  • Practical experience in AIOps, observability, monitoring and operational analytics across diverse production environments.
  • Ability to lead and influence in a matrixed global organisation, including cross-functional CTO engineering, infrastructure and application teams.
  • Experience in Application Performance Management and modern telemetry practices across metrics, logs, traces and event correlation.
  • Evidence of troubleshooting complex incidents in high-pressure environments in an evidence-driven, systemic and outcome-focused manner.
  • Demonstrable experience improving change quality, operational resilience and service reliability across multiple teams.
  • Sound knowledge of platform management and operations, ideally with practical adoption of SRE principles and reliability engineering methods.
  • Experience working with TPEMs, delivery stakeholders and vendor account managers to manage dependencies and deliver outcomes.
  • Subject matter experience across multiple technology classes and operating environments in financial services, including core banking, cloud-native, multi-CSP, hybrid and data centre estates.
  • Strong hands-on programming capability, ideally in Python or Java for automation, API integration, data handling and AI-assisted operational tooling; experience with JavaScript, TypeScript or enterprise scripting languages would be advantageous.
  • Practical experience of secure API and tool integration with clear controls for identity, access, auditability and production safety.
  • Ability to communicate complex technical issues clearly to all levels of the organisation, from engineers to CIO and business stakeholders.
  • Strong written and spoken English, with confidence operating in a multi-cultural, diverse and inclusive organisation.
  • Demonstrate decisive judgement, strong ownership and the ability to deliver high-quality outcomes on time.
  • Personal behaviours aligned to the role: ego-less collaboration, curiosity, tenacity, precision, and resilience under pressure.

The employment is subject to Mandatory Reference Checking Scheme (MRCS) as per regulatory requirement. For details, please refer to (Mandatory Reference Checking Scheme Phase 2 | The Hong Kong Association of Banks).

Opening up a world of opportunity www.hsbc.com/careers

HSBC is committed to building a culture where all employees are valued, respected and opinions count. We take pride in providing a workplace that fosters continuous professional development, flexible working and opportunities to grow within an inclusive and diverse environment. Personal data held by the Bank relating to employment applications will be used in accordance with our Privacy Statement, which is available on our website.

Issued by The Hongkong and Shanghai Banking Corporation Limited.

https://www.youtube.com/embed/QmZ7Un5gR8c?si=LCa6slfBqlUlxUE-

Similar roles