Sophos

Manager, Software Engineering

Sophos United States · $153K–$255K/yr

Software Development · 5,001-10,000 employees

5 d ago
Remote Principal (10+ yrs) Full-time United States
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

Lead a distributed team of engineers to improve the reliability, scalability, and operational maturity of cloud services. Drive incident response, observability, automation, and cross-functional collaboration to ensure high-availability systems.

What they look for

Site Reliability Engineering Team Leadership Cloud Operations AWS Kubernetes CI/CD Infrastructure as Code Observability Incident Management Automation Linux Terraform Ansible Python Agile Security Practices

Requirements

Requires extensive experience in site reliability engineering, DevOps, and cloud infrastructure management, particularly with AWS and Kubernetes. Strong leadership skills and the ability to manage technical delivery, security compliance, and stakeholder communication are essential.

Benefits

Bonus eligibility Comprehensive benefits package Wellbeing days Wellbeing webinars Volunteer days Diversity and inclusion networks

Full description

About Us

Sophos is a cybersecurity leader defending 600,000 organizations globally with an AI-driven platform and expert-led services. Sophos meets organizations wherever they are in their security maturity and grows with them to defeat cyberattacks. Its solutions combine machine learning, automation, and real-time threat intelligence with frontline human expertise from Sophos X-Ops to deliver advanced, 24/7 threat monitoring, detection, and response.

Sophos offers industry-leading managed detection and response (MDR) alongside a comprehensive portfolio of cybersecurity technologies — including endpoint, network, email, and cloud security, extended detection and response (XDR), identity threat detection and response (ITDR), and next-gen SIEM. Together with expert advisory services, these capabilities help organizations proactively reduce risk and respond faster, with the visibility and scalability needed to stay ahead of evolving threats.

Sophos goes to market with a global partner ecosystem, including Managed Service Providers (MSPs), Managed Security Service Providers (MSSPs), resellers and distributors, marketplace integrations, and cyber risk partners, giving organizations the flexibility to choose trusted relationships when securing their business. Sophos is headquartered in Oxford, U.K. More information is available at www.sophos.com.

Role Summary

We are seeking an experienced Manager, Software Engineering (SRE) to lead a team focused on improving the reliability, scalability, and operational maturity of Sophos cloud services and engineering delivery systems. In this role, you will lead a team distributed across the U.S. and Canada. The team is responsible for production reliability, operational readiness, incident response, escalation management, observability, automation, release reliability, and continuous improvement across cloud-based platforms.

This role requires a leader with a strong background spanning site reliability engineering, platform engineering, DevOps practices, release engineering, infrastructure management, and cloud operations. You should have enough technical depth to guide discussions around AWS, Kubernetes/EKS, CI/CD, Linux, infrastructure as code, observability, cost optimization, security practices, compliance needs, and automation, while primarily focusing on team leadership, execution, stakeholder alignment, prioritization, and improving how engineering teams deliver and operate services at scale.

\n

What You Will Do Leadership and Team Management

  • Lead, coach, and develop a team of engineers focused on site reliability, platform operations, release engineering, and automation. Set clear priorities, expectations, and delivery goals aligned to business and engineering needs.
  • Support hiring, onboarding, performance management, and career development for team members.
  • Foster a culture of ownership, collaboration, continuous improvement, and operational excellence.

SRE Strategy and Operational Maturity

  • Partner with senior leadership to define and execute the roadmap for SRE and operational improvements.
  • Drive improvements in reliability, availability, scalability, deployment confidence, and production readiness.
  • Drive operational efficiency and cost optimization practices across cloud platforms, balancing reliability, performance, and responsible cloud spend.
  • Help establish best practices for incident response, escalation management, observability, runbooks, post-incident reviews, and toil reduction.
  • Promote the use of reliability metrics, operational health indicators, and continuous improvement practices.

Platform, Release, and Automation Enablement

  • Partner with engineering, architecture, security, release, and operations teams to improve shared tooling and delivery workflows.
  • Support improvements across CI/CD pipelines, release automation, deployment reliability, and environment stability. Guide team efforts involving cloud infrastructure, Kubernetes/EKS, Linux systems, Terraform/IaC, Ansible, and automation tooling.
  • Encourage responsible use of approved AI tools to improve troubleshooting, documentation, automation, and engineering productivity.

Incident Management and Reliability

  • Ensure effective incident and escalation management practices are in place for timely response, communication, resolution, and follow-up.
  • Assess operational situations quickly, prioritize response activities, and balance customer impact, business risk, reliability, security, and delivery needs.
  • Drive root cause analysis and long-term corrective actions to reduce repeat incidents.
  • Improve observability, alerting, monitoring, and operational visibility across services and platforms.
  • Support capacity planning and operational readiness for scalable, high-availability systems.

Security, Compliance, and Governance

  • Partner with security, AppSec, compliance, and engineering teams to ensure operational practices support secure and reliable service delivery.
  • Help manage CI/CD security gates, vulnerability remediation workflows, audit readiness, and compliance-related operational controls.
  • Support vendor management for SRE, platform, observability, security, and cloud tooling, including renewals, evaluations, usage reviews, and service escalations.
  • Ensure team documentation, runbooks, operational evidence, and process controls are maintained to support audits and internal reviews.

Planning and Delivery

  • Lead sprint planning, backlog management, prioritization, and delivery tracking for SRE and platform initiatives.
  • Balance competing priorities across reliability work, operational escalations, security needs, technical debt, roadmap delivery, and team capacity.
  • Participate in quarterly planning, roadmap development, and cross-team dependency management.
  • Work with engineering, product, architecture, and security stakeholders to align reliability and operational priorities with business needs.
  • Use Agile/Scrum practices to improve execution, transparency, predictability, and continuous improvement across the team.

Collaboration and Communication

  • Work closely with engineering teams to understand service needs and ensure SRE priorities align with product and platform goals.
  • Lead technical and operational discussions with engineering leaders, product stakeholders, and executive audiences, including Directors, VPs, and C-level staff.
  • Present reliability updates, operational risks, roadmap progress, incident trends, and investment needs in a clear, concise, and audience-appropriate way.
  • Facilitate meetings, drive decisions, manage follow-ups, and ensure cross-functional conversations result in clear ownership and action.
  • Communicate status, risks, operational trends, and improvement opportunities to technical and non-technical stakeholders.

What You Will Bring

  • Extensive experience in site reliability engineering, DevOps practices, platform engineering, infrastructure management, operational efficiency, cloud infrastructure, or production operations.
  • Proven experience leading and managing engineers, fostering a culture of collaboration, accountability, ownership, and continuous improvement.
  • Strong understanding of reliability engineering, operational excellence, incident management, escalation management, observability, production support, and scalable system operations.
  • Strong working knowledge of cloud technologies, especially AWS; Azure experience is a plus.
  • Familiarity with operating and supporting containerized platform environments, including Kubernetes/EKS, Docker, Helm, and Linux-based systems.
  • Familiarity with infrastructure as code, configuration management, and automation practices using tools such as Terraform and Ansible.
  • Understanding of CI/CD pipelines, release automation, deployment practices, and developer tooling such as Jenkins, GitHub Actions, Artifactory, or similar tools.
  • Proficiency with scripting and automation concepts using languages such as Python, Bash, or Groovy.
  • Experience with monitoring, logging, and observability tools such as ELK stack, Prometheus, Grafana, CloudWatch, or similar platforms.
  • Working knowledge of security, application security, vulnerability management, compliance controls, audit readiness, and vendor management in a cloud or SaaS environment.
  • Familiarity with Agile/Scrum practices, sprint planning, backlog management, quarterly planning, roadmap development, and cross-team dependency management.
  • Excellent problem-solving and analytical skills, with the ability to assess situations quickly, prioritize work, manage escalations, and balance competing technical and business priorities.
  • Strong presentation and facilitation skills, with the ability to communicate clearly and confidently with technical teams, senior leaders, executives, vendors, and cross-functional stakeholders.
  • Ability to guide technical discussions, evaluate trade-offs, balance reliability, security, compliance, delivery speed, and operational efficiency, and help teams make sound engineering decisions.
  • Experience improving operational processes, reducing manual toil, mentoring engineers, managing delivery, and building healthy, accountable teams.
  • Experience with Java-based services, build systems, or code review workflows is a plus.
  • Comfort using or encouraging AI-assisted engineering tools responsibly within company guidelines.
  • Bachelor’s degree in Computer Science, Information Technology, or a related field, or equivalent practical experience. A Master’s degree is a plus.

\n In the United States, the base salary for this role ranges from $153,000 to $255,000. In addition to base salary, we offer additional compensation including bonus eligibility and a comprehensive benefits package.  A candidate’s specific pay within this range will depend on a variety of factors, including job-related skills, training, location, experience, relevant education, certifications, and other business and organizational needs.

#li-remote

#b2

#li-ND2

Ready to Join Us?

At Sophos, we believe in the power of diverse perspectives to fuel innovation. Research shows that candidates sometimes hesitate to apply if they don't check every box in a job description. We challenge that notion. Your unique experiences and skills might be exactly what we need to enhance our team. Don't let a checklist hold you back – we encourage you to apply.

What's Great About Sophos?

· Sophos operates a remote-first working model, making remote work the primary option for most employees. However, some roles may necessitate a hybrid approach. While we are a remote first organization, applicants must have legal authorization to work in the jurisdiction where the position is posted, without requiring employer sponsorship.

· Our people – we innovate and create, all of which are accompanied by a great sense of fun and team spirit

· Employee-led diversity and inclusion networks that build community and provide education and advocacy

· Annual charity and fundraising initiatives and volunteer days for employees to support local communities

· Global employee sustainability initiatives to reduce our environmental footprint

· Global fitness and trivia competitions to keep our bodies and minds sharp

· Global wellbeing days for employees to relax and recharge

· Monthly wellbeing webinars and training to support employee health and wellbeing

Our Commitment To You

We’re proud of the diverse and inclusive environment we have at Sophos, and we’re committed to ensuring equality of opportunity. We believe that diversity, combined with excellence, builds a better Sophos, so we encourage applicants who can contribute to the diversity of our team. All applicants will be treated in a fair and equal manner and in accordance with the law regardless of gender, sex, gender reassignment, marital status, race, religion or belief, color, age, military veteran status, disability, pregnancy, maternity or sexual orientation. We want to give you every opportunity to show us your best self, so if there are any adjustments we could make to the recruitment and selection process to support you, please let us know.

Data Protection

If you choose to explore an opportunity, and subsequently share your CV or other personal details with Sophos, these details will be held by Sophos for 12 months in accordance with our Privacy Policy and used by our recruitment team to contact you regarding this or other relevant opportunities at Sophos. If you would like Sophos to delete or update your details at any time, please follow the steps set out in the Privacy Policy describing your individual rights. For more information on Sophos’ data protection practices, please consult our Privacy Policy Cybersecurity as a Service Delivered | Sophos