Aviva

Senior Manager, Site Reliability & Infrastructure Engineering

Aviva · Markham, Ontario, Canada · CA$125K–CA$175K/yr

Insurance · 10,001+ employees

19 h ago
Principal (10+ yrs) Full-time Canada
Log in to apply, save this posting, or score it against your profile with AI.

About the role

The Senior Manager will lead the evolution toward an engineering-led reliability and resilience capability across critical applications and platforms. They will manage a team of engineers to mature SRE practices, oversee technology resilience, and ensure audit-ready operational performance.

What they look for

Site Reliability Engineering Infrastructure Engineering Cloud Computing Technology Resilience Disaster Recovery Cyber Recovery Observability Incident Response Automation AWS Rubrik Dynatrace Leadership Risk Management Vendor Management Problem Management

Requirements

Candidates must have 10+ years of experience in SRE, infrastructure, or related disciplines, including 5+ years of leadership experience. A bachelor's degree in Computer Science, Engineering, or Information Systems is required.

Benefits

Annual bonus Retirement savings Share plan Health benefits Personal wellness Volunteer opportunities Career development Vacation package Employee discount Hybrid flexible work model

Full description

Individually we are people, but together we are Aviva. Individually these are just words, but together they are our Values – Care, Commitment, Community, and Confidence.

At Aviva Canada, we put people first, our employees, our customers, and our communities. We’re proud of a culture built on care, inclusion, and collaboration, where your voice matters and your growth is supported. We’re not just about insurance; we’re about making a real difference by protecting what matters most.

The Senior Manager, Site Reliability & Infrastructure Engineering will lead Aviva Canada’s evolution toward an engineering-led reliability and resilience capability across critical applications, platforms and services. This role will work across on-premises, AWS, SaaS/vendor-hosted and hybrid application environments, partnering with Engineering Operations, Cloud & Platform Engineering, Application Engineering, Cybersecurity, Operational Resilience, Business Continuity, Enterprise Architecture, Vendor Management and key partners including AWS, DXC, CGI, Snowflake and other SaaS providers.

The focus is to shift Aviva from periodic recovery and resilience testing to proactive, measurable reliability engineering. This uses observability platforms like Dynatrace, incident response tools such as PagerDuty or equivalent, and modern recovery platforms like Rubrik to improve service availability, operational insight, recovery readiness, and customer outcomes.

What you’ll do

Site Reliability Engineering

  • Establish and mature SRE practices across critical applications and platforms, including SLIs, SLOs, SLAs, service health indicators, post-incident reviews and continuous reliability improvement.
  • Partner with application, platform, infrastructure and vendor teams to embed reliability requirements into design, development, operational acceptance, release and production support processes.
  • Drive improvements in incident response, problem management, root cause analysis, alert quality, critical issue workflows, service health reporting and reduction of repeat incidents.
  • Improve Infrastructure Ops via automation, self-healing patterns, runbook improvements, AI-assisted operations and repeatable engineering practices.
  • Use observability insights to improve infrastructure availability, performance, capacity planning, & operational readiness.

Technology Resilience, Backup & Cyber Recovery

  • Maintain and evolve Aviva Canada’s technology resilience capability across on-premises platforms, AWS, critical applications, integration services and SaaS/vendor-hosted platforms.
  • Ensure resilience planning validates end-to-end recoverability, including application, data, integration, identity, network, platform, cloud and vendor dependencies.
  • Support annual BCP, DR and technology resilience testing, including RTO/RPO validation, dependency mapping, recovery sequencing, test evidence, gap management and remediation tracking.
  • Own and mature backup and cyber recovery practices, including hands-on use of Rubrik or a comparable enterprise data protection platform for backup policy management, immutability, restore validation, coordinating recovery procedures, reporting, evidence capture and operational support for critical workloads.
  • Partner with Cybersecurity on ransomware and destructive cyber event readiness, including clean recovery, backup integrity validation, isolated recovery environments, cyber recovery vault operations and restoration of critical services.
  • Ensure resilience and recovery requirements are embedded into architecture, cloud migration, vendor onboarding, operational readiness, change delivery and service governance.

Leadership

  • Lead, manage and coach group of engineers working across reliability, observability, production support, technology resilience and recovery practices.
  • Build a clear operating model for SRE and technology resilience ownership across application teams, infrastructure teams, cloud/platform teams, cybersecurity, operational resilience and third-party providers.
  • Define reliability and resilience reporting for senior leaders, including service health, SLO performance, incident trends, MTTR, alert quality, recovery readiness, open risks and remediation status.
  • Partner with Operational Resilience, Business Continuity, Technology Risk and Cybersecurity teams to ensure evidence, controls and remediation plans are audit-ready and aligned to regulatory expectations.
  • Influence engineering and operations teams to adopt reliability-by-design, automation-first and evidence-based ways of working.
  • Serve as a designated point of contact for material risks relating to service reliability, monitoring gaps, application recoverability, backup coverage, cyber recovery readiness, RTO/RPO gaps and vendor resilience ambiguity.
  • Partner with key infrastructure managed service providers to continuously improve service quality, strengthen operational performance, hold vendors accountable for meeting agreed service levels, outcomes, remediation commitments and continuous improvement targets.

What you’ll bring

  • 10+ years of technology experience across Site Reliability Engineering, Infrastructure Engineering, Platform Engineering, Cloud, Technology resilience, Disaster recovery, Cyber recovery or related disciplines.
  • 5+ years of experience leading & managing technical teams, with the ability to mentor engineers, set direction, manage priorities and high visible projects.
  • Experience with enterprise observability tooling such as Dynatrace, Grafana, Datadog or equivalent platforms.
  • Hands-on experience operating Rubrik or a comparable enterprise backup and cyber recovery platform, including backup policy configuration, immutable backup concepts, restore testing, recovery workflow execution, access controls, reporting, evidence capture and support for critical workloads.
  • Experience supporting distributed, business-critical applications across hybrid environments, including on-premises infrastructure, cloud platforms, APIs, middleware, databases, containers and SaaS/vendor-hosted systems.
  • Working knowledge of AWS or comparable cloud platforms, including cloud resilience, infrastructure as code, monitoring, logging, backup/recovery, identity, network dependencies and landing zone operating models.
  • Demonstrated experience with incident, problem, change, and release management practices across complex and/or highly regulated environments.
  • Knowledge of cyber recovery concepts, including ransomware recovery, clean-room recovery, isolated recovery environments, backup integrity validation and cyber incident recovery planning.
  • Strong written and verbal communication skills, including demonstrating proficiency in producing clear executive reporting, operational dashboards and audit-ready evidence.
  • Bachelor’s degree in Computer Science, Engineering, Information Systems or equivalent.
  • P&C insurance domain experience, Guidewire/Snowflake exposure, AWS/Rubrik/Dynatrace certifications would be considered assets.

What you’ll get

  • The salary band for this position ranges from $125,000 – $175,000. Please note that individual salary is determined by factors such as job-related knowledge, skills and experience, as well as internal equity.
  • Compelling rewards package including base compensation, eligibility for annual bonus, retirement savings, share plan, health benefits, personal wellness, and volunteer opportunities.
  • Outstanding Career Development opportunities.
  • We’ll support your professional development education.
  • Competitive vacation package with the option to purchase 5 extra days off per year.
  • Employee driven programs focused on gender, LGBTQ+, origins, diversity, and inclusion.
  • Corporate wellness programs to support our employees’ physical and mental health.
  • Employee discount on home and auto insurance (where applicable).
  • Hybrid flexible work model.

This job advertisement is for a new vacancy which has been posted both internally & externally.

Aviva Canada may use AI (Artificial Intelligence) tools to assist us throughout the recruitment process to screen, assess or select applicants for a position.

Aviva Canada welcomes applications from all qualified individuals and has a process in place to provide accommodations for persons with disabilities at all stages of the hiring process and during employment. If you require an accommodation during the interview or hiring process, please contact your Aviva Talent Acquisition Partner so that an appropriate accommodation can be arranged.

#LI-PS1 #LI-Hybrid