OneAdvanced

Senior Engineering Manager - SRE

OneAdvanced Bangalore North, Karnataka, India

Software Development · 1,001-5,000 employees

14 h ago
sre Principal (10+ yrs) Other India
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

Lead the Site Reliability Engineering function by defining the monitoring and observability roadmap across a diverse product portfolio. Drive innovation by building and deploying AI agents to automate incident detection, alerting, and resolution.

What they look for

Site reliability engineering Observability Platform engineering AI agents Cloud-native technologies AWS Kubernetes Infrastructure as code CI/CD Distributed systems Incident management Team leadership Monitoring Automation Operational excellence Stakeholder management

Requirements

Requires 10+ years of experience in SRE or platform engineering with a proven track record of leading large teams. Candidates must possess hands-on expertise in building observability platforms and deploying AI-driven automation for incident management.

Benefits

Annual leave Public holidays Employee assistance programme Development programmes Online learning platform Life insurance Personal accident insurance Performance bonus

Full description

Join OneAdvanced

About the Role

OneAdvanced's Site Reliability Engineering function is responsible for how well we can see into our own systems, and how fast we can act on what we see. As Senior Manager, Site Reliability Engineering, you'll own the monitoring and observability roadmap across our full product portfolio — spanning Health, Legal, Education, Workforce Management, and other sectors — and lead a team of 15-20 SRE engineers to deliver it.

You'll also push the function further: building and deploying AI agents that handle monitoring, alerting, and incident resolution directly — reducing how much of this work depends on a person watching a dashboard, and moving us toward a more autonomous operating model.

Mandate: To lead OneAdvanced's Site Reliability Engineering function — building the monitoring and observability roadmap that gives every product across Health, Legal, Education, Workforce Management, and other sectors real visibility into its own health, and increasingly uses AI agents to detect, alert on, and resolve incidents before they need a person.

What You Will Do

Key Responsibilities:

Monitoring & Observability Roadmap

  • Own and drive the monitoring and observability roadmap across the product portfolio — Health, Legal, Education, Workforce Management, and other sectors.
  • Set the standard for what “well-monitored” means for a product, and hold the portfolio to it.
  • Evaluate, select, and evolve the monitoring toolset to match the scale and complexity of the estate.
  • Lead, coach and inspire a high-performing team of Site Reliability Engineers.
  • Shape and deliver our Site Reliability Engineering roadmap alongside the Head of Platform.
  • Champion modern engineering practices including SLIs, SLOs, error budgets, observability and automation.
  • Improve the reliability, scalability and performance of our cloud platforms and digital services.
  • Partner with Engineering, Security, Data and Product teams to embed operational excellence from design through to production.
  • Drive the adoption of our observability platform, helping teams gain deeper insight into the health and performance of their services.
  • Lead incident learning, continuous improvement and automation initiatives that reduce operational toil.
  • Provide technical leadership across AWS, Kubernetes, Infrastructure as Code, CI/CD and distributed system

AI-Driven Monitoring & Incident Resolution

  • Design, build, and deploy AI agents that handle monitoring, alerting, and first-line incident resolution.
  • Identify where agentic automation can safely replace manual triage, and where human judgment still needs to stay in the loop.
  • Continuously improve agent accuracy and trust based on real incident outcomes.

Incident Management

  • Partner closely with Major Incident Management to ensure observability data drives faster detection and diagnosis during live incidents.
  • Drive root cause analysis for monitoring or alerting gaps that contributed to incident impact or delay.
  • Ensure post-incident learnings translate into concrete monitoring and alerting improvements.

Team Leadership

  • Lead, grow, and mentor a team of 10–15 SRE engineers, building deep observability and automation expertise.
  • Own hiring, performance management, and career development for the team.
  • Build a culture of ownership, curiosity, and continuous improvement.

Cross-Functional Partnership

  • Work closely with Engineering, Services, and Customer Success teams to ensure monitoring reflects what actually matters to product reliability and customer experience.
  • Represent SRE & Observability in cross-functional planning and governance forums.
  • Act as an escalation point for observability and monitoring gaps raised by any stakeholder team.

#LI-PB1

What You Will Have

  • 10+ years of experience in Site Reliability Engineering, observability, or platform engineering, including people leadership experience.
  • A proven track record leading teams of 15+ engineers.
  • Hands-on experience building and operating monitoring and observability platforms at scale.
  • Practical experience building or deploying AI agents for monitoring, alerting, or incident resolution — not just familiarity with the concept.
  • Experience partnering with incident management functions to improve detection and response.
  • Strong communication skills, comfortable working across engineering, services, and customer-facing teams.
  • Experience leading Site Reliability Engineering, Platform Engineering, DevOps or Cloud Infrastructure teams.
  • Strong technical expertise in cloud-native technologies, ideally within AWS, Azure and Private cloud platforms.
  • Experience operating and improving large-scale distributed systems.
  • A passion for coaching, mentoring and helping engineers thrive.
  • Experience driving observability, automation and operational excellence.
  • The ability to build strong relationships and influence stakeholders across Engineering, Product, Security and Data.
  • A pragmatic approach to balancing reliability, innovation and customer impact.

It would be great if you also had

  • Experience with monitoring and observability tools such as Grafana, LogicMonitor, or equivalent.
  • Experience with ServiceNow and Jira for incident and delivery tracking.
  • Hands-on experience with AI-assisted engineering tools (e.g. Claude Code, GitHub Copilot) beyond monitoring use cases.
  • Experience operating across multi-sector or multi-product portfolios.
  • Relevant certifications in observability platforms or cloud providers (AWS, Azure)

What We Do For You

  • Wellbeing focused – Our people are our greatest assets, and ensuring everyone feels their best self to come to work is integral.
  • Annual Leave – 20 days of annual leave, plus public holidays 
  • Employee Assistance Programme – Free advice, support, and confidential counselling available 24/7.
  • Personal Growth - Regardless of where you are at in your career, we’re committed to enabling your growth personally and professionally
  • Development Programmes – From Future Managers to Leadership Training, our development programmes help you get where you need to go
  • Online Learning Platform: SkillsHub! - Learning at your fingertips, anytime from anywhere. You can access our online library with relevant content for your career growth.
  • Life Insurance - 3x annual salary
  • Personal Accident Insurance - providing cover in the event of serious injury/illness.
  • Performance Bonus – Our Group-wide bonus scheme enables you to reap the rewards of your success

Who We Are

At OneAdvanced, we are at the forefront of delivering sector-focused technology solutions that simplify complexity, drive meaningful progress, and help build a fairer, more inclusive society.

We’re much more than a software company. We deliver SaaS workflow applications and IT services that power organisations across Education, Government, Healthcare, Legal, Manufacturing, Housing, Retail, and more.

OneAdvanced is one of the UK’s largest business software and services companies. Based in Birmingham (The Mailbox), operating across the UK, Ireland, India, and Australia.

Our secure, scalable platform, including OneAdvanced AI, our private AI service for UK organisations, powers connectivity and innovation across critical sectors. Alongside our software are our IT services, including hosting, managed services, and application modernisation.

We strive to create an inclusive workplace that drives innovation and collaboration, championing diverse perspectives and ideas. Our Environmental, Social and Governance (ESG) strategy is embedded in everything we do, guiding us to create meaningful impact for our people, our customers and the planet.

Join us and become part of a team that’s powering the world of work and making a real difference.

Learn more at www.oneadvanced.com

Similar roles