Barclays

Senior Caching SRE

Barclays Knutsford, England, United Kingdom

Banking · 10,001+ employees

20 h ago Closes in 6d
sre Senior (5-10 yrs) Full-time United Kingdom
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

The Senior Caching SRE will ensure the reliability, availability, and scalability of caching platforms through proactive monitoring, maintenance, and capacity planning. They will also develop automation tools and collaborate with development teams to integrate reliability best practices into the software development lifecycle.

What they look for

Redis Kubernetes OpenShift Python Site Reliability Engineering Incident Management Observability Automation Performance Tuning Capacity Planning Cluster Management Service Reliability System Resilience Kafka AIOP

Requirements

Candidates must have strong experience operating and troubleshooting Redis at scale and deploying applications on Kubernetes or OpenShift. Proficiency in Python for automation and a solid background in SRE principles and incident management are also required.

Full description

Job Description

Purpose of the role

To apply software engineering techniques, automation, and best practices in incident response, to ensure the reliability, availability, and scalability of the systems, platforms, and technology through them. 

Accountabilities

  • Availability, performance, and scalability of systems and services through proactive monitoring, maintenance, and capacity planning.
  • Resolution, analysis and response to system outages and disruptions, and implement measures to prevent similar incidents from recurring.
  • Development of tools and scripts to automate operational processes, reducing manual workload, increasing efficiency, and improving system resilience.
  • Monitoring and optimisation of system performance and resource usage, identify and address bottlenecks, and implement best practices for performance tuning.
  • Collaboration with development teams to integrate best practices for reliability, scalability, and performance into the software development lifecycle, and work closely with other teams to ensure smooth and efficient operations.
  • Stay informed of industry technology trends and innovations, and actively contribute to the organization's technology communities to foster a culture of technical excellence and growth.

Vice President Expectations

  • To contribute or set strategy, drive requirements and make recommendations for change. Plan resources, budgets, and policies; manage and maintain policies/ processes; deliver continuous improvements and escalate breaches of policies/procedures..
  • If managing a team, they define jobs and responsibilities, planning for the department’s future needs and operations, counselling employees on performance and contributing to employee pay decisions/changes. They may also lead a number of specialists to influence the operations of a department, in alignment with strategic as well as tactical priorities, while balancing short and long term goals and ensuring that budgets and schedules meet corporate requirements..
  • If the position has leadership responsibilities, People Leaders are expected to demonstrate a clear set of leadership behaviours to create an environment for colleagues to thrive and deliver to a consistently excellent standard. The four LEAD behaviours are: L – Listen and be authentic, E – Energise and inspire, A – Align across the enterprise, D – Develop others..
  • OR for an individual contributor, they will be a subject matter expert within own discipline and will guide technical direction. They will lead collaborative, multi-year assignments and guide team members through structured assignments, identify the need for the inclusion of other areas of specialisation to complete assignments. They will train, guide and coach less experienced specialists and provide information affecting long term profits, organisational risks and strategic decisions..
  • Advise key stakeholders, including functional leadership teams and senior management on functional and cross functional areas of impact and alignment.
  • Manage and mitigate risks through assessment, in support of the control and governance agenda.
  • Demonstrate leadership and accountability for managing risk and strengthening controls in relation to the work your team does.
  • Demonstrate comprehensive understanding of the organisation functions to contribute to achieving the goals of the business.
  • Collaborate with other areas of work, for business aligned support areas to keep up to speed with business activity and the business strategies.
  • Create solutions based on sophisticated analytical thought comparing and selecting complex alternatives. In-depth analysis with interpretative thinking will be required to define problems and develop innovative solutions.
  • Adopt and include the outcomes of extensive research in problem solving processes.
  • Seek out, build and maintain trusting relationships and partnerships with internal and external stakeholders in order to accomplish key business objectives, using influencing and negotiating skills to achieve outcomes.

All colleagues will be expected to demonstrate the Barclays Values of Respect, Integrity, Service, Excellence and Stewardship – our moral compass, helping us do what we believe is right. They will also be expected to demonstrate the Barclays Mindset – to Empower, Challenge and Drive – the operating manual for how we behave.

Join Barclays as a Senior Caching SRE and play a pivotal role in delivering highly reliable, scalable, and resilient caching platforms that support critical business services. You will drive operational excellence, enhance system performance, and automate platform capabilities while ensuring exceptional service availability across a complex enterprise environment.

To be successful as a Senior Caching SRE, you should have experience with:

  • Strong experience operating and troubleshooting Redis at scale, including Cluster, Replication, HA, Backup/Recovery, and Performance Tuning – ensuring high availability, performance optimisation, and rapid issue resolution across mission-critical environments
  • Practical experience deploying and operating applications on Kubernetes/OpenShift, including deployments, upgrades, scaling, and troubleshooting in a production environment – ensuring stable and scalable platform operations
  • Strong SRE background with expertise in service reliability, incident management, observability, automation, and operational excellence – driving platform stability, resilience, and continuous service improvement
  • Strong Python programming skills for platform automation, self-service tooling, operational workflows, and service lifecycle management – enabling scalable automation and increased operational efficiency

Some other highly valued skills may include:

  • Working experience in financial domain
  • Experience with Kafka based event-driven architecture
  • Experience with AIOP process and methods

You may be assessed on the key critical skills relevant for success in role, such as risk and controls, change and transformation, business acumen strategic thinking and digital and technology, as well as job-specific technical skills.

This role is based in Knutsford

Similar roles