Lloyds Banking Group

Senior Site Reliability Engineer

Lloyds Banking Group Hyderabad, Telangana, India

Financial Services · 10,001+ employees

12 h ago Closes today
sre Principal (10+ yrs) Full-time India
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

The Senior Site Reliability Engineer will design and implement observability solutions to improve service reliability and reduce operational toil. They will also collaborate with engineering teams to embed best practices and mentor junior engineers to promote technical excellence.

What they look for

Site Reliability Engineering Observability Monitoring Automation Incident Management Infrastructure as Code CI/CD Python Java JavaScript Cloud technologies Kubernetes Azure Google Cloud Platform AWS Agile

Requirements

Candidates must have 9-15 years of experience with strong knowledge of SRE principles, observability tools, and cloud platforms. Proficiency in scripting languages and experience with incident management and automation in complex production environments is required.

Benefits

Flexible working Hybrid working

Full description

End Date

Friday 11 September 2026

We Support Flexible Working – Click here for more information on flexible working options

Flexible Working Options

Hybrid Working

Job Description Summary

Our Site Reliability Engineering (SRE) team plays a key role in improving the reliability, resilience and operational excellence of the Analytics & AI Platform.

As a Site Reliability Engineer, you’ll work closely with Product, Engineering and Production Support teams to improve service reliability, reduce operational risk and embed Site Reliability Engineering principles across our platforms. You’ll use engineering, automation and observability to improve service resilience, reduce operational toil and support the delivery of highly available services.

As an experienced engineer, you’ll have opportunities to mentor colleagues, lead technical initiatives and, where appropriate, take on line management responsibilities to support the growth and development of engineers within the team

Job Description

  • Role: Senior Site Reliability Engineer

Experience: 9-15 years Location: Hyderabad Job Type: Full Time  

  • What you’ll do 
  • Design, implement and continuously improve monitoring, alerting and observability solutions that enable reliable production services. 
  • Define, measure and continuously improve Service Level Indicators (SLIs), Service Level Objectives (SLOs) and Error Budgets to drive service reliability. 
  • Work collaboratively with Production Support and engineering teams to investigate complex production incidents and support service restoration. 
  • Support Post Incident Reviews (PIRs) and Problem Management activities by identifying reliability improvements and driving preventative engineering actions. 
  • Reduce operational toil through automation, tooling and engineering improvements, enabling teams to focus on higher-value engineering work. 
  • Work closely with Production Support teams to develop and continuously improve operational runbooks, support processes and service readiness. 
  • Partner with engineering teams to embed reliability, resilience and operational best practices throughout the software development lifecycle. 
  • Collaborate with Product and Engineering teams to improve service operability, ensuring applications are designed with appropriate monitoring, alerting, diagnostics and operational documentation before entering production support. 
  • Identify reliability risks and proactively deliver engineering improvements that reduce incidents and improve service resilience. 
  • Support the onboarding of new applications into the SRE operating model by ensuring agreed reliability, observability and operational standards are achieved. 
  • Coach and mentor engineers, sharing knowledge and promoting engineering excellence across the platform. 
  • Contribute to the evolution of SRE standards, tooling and engineering practices across the Analytics & AI Platform. 
  • Where appropriate, provide line management, coaching and performance development for engineers within the team. 

  

What you’ll need 

 

Essential 

  • Strong understanding of Site Reliability Engineering principles and practices. 
  • Experience designing and implementing observability solutions, including monitoring, logging and alerting. 
  • Experience defining and improving Service Level Indicators (SLIs), Service Level Objectives (SLOs) and Error Budgets. 
  • Strong troubleshooting and technical investigation skills across complex production environments. 
  • Experience supporting incident management, Post Incident Reviews and Problem Management activities. 
  • Experience identifying operational toil and delivering automation to improve reliability and efficiency. 
  • Experience using Infrastructure as Code and CI/CD tooling. 
  • Strong scripting or programming skills in one or more languages such as Python, Java, JavaScript, PowerShell or Bash. 
  • Experience working with cloud technologies and modern application platforms. 
  • Strong understanding of cloud security, networking and operational resilience. 
  • Excellent stakeholder management, communication and collaboration skills. 
  • Ability to work effectively across Product, Engineering and Production Support teams. 

 

Desirable 

  • Experience with Kubernetes and containerised platforms. 
  • Experience with Azure, Google Cloud Platform or AWS. 
  • Experience using observability platforms such as Dynatrace, Grafana, Prometheus, ELK or Splunk. 
  • Experience working within Financial Services or another regulated industry. 
  • Experience using Jira, Confluence and Agile delivery practices. 
  • Experience mentoring or coaching engineers. 
  • Previous line management experience or a desire to develop people leadership capability. 
  • Relevant Cloud, DevOps or SRE certifications

Similar roles