Manager, Site Reliability Engineering
Open Text (Philippines), Inc. Makati, National Capital District, Philippines
Software Development · 10,001+ employees
About the role
The role involves leading day-to-day cloud operations on AWS to ensure system stability, security, and performance. You will manage incident response, proactive monitoring, and infrastructure maintenance while mentoring operational staff.
What they look for
Requirements
Candidates must have at least 5 years of experience in cloud operations or systems administration with strong expertise in AWS. A tertiary qualification in Computer Science or a related field is preferred, along with proficiency in infrastructure automation and monitoring tools.
Full description
Role Overview
We are seeking an experienced Cloud Operations Manager to lead and mature the operational management of a cloud environment hosted on AWS. This is a hands-on leadership role responsible for the stability, security, performance, and continuous improvement of both production and non-production environments.
You will lead a Cloud/Production Services function that ensures applications and infrastructure are deployed, operated, and optimized in alignment with architectural standards, security controls, and operational best practices. Working closely with software engineering, security, infrastructure, and customer support teams, you will drive operational excellence while meeting compliance and availability requirements.
This role includes ownership of 24x7x365 operational readiness, including on-call rotations, incident response, and proactive monitoring.
Key Responsibilities
- Lead day-to-day cloud operations for an AWS environment, ensuring compliance with security, availability, and governance requirements
- Own the health, performance, and capacity of production and non-production environments
- Provide hands-on technical leadership across cloud infrastructure, applications, and monitoring
- Ensure application deployments and operational practices align with overall architecture and security standards
- Oversee maintenance activities including upgrades, patching, backups, and recovery testing
- Drive proactive monitoring, alerting, and incident management to maintain high availability and reliability
- Act as an escalation point for complex technical issues and production incidents
- Manage and participate in an on-call rotation supporting 24x7x365 operations
- Identify operational inefficiencies and continuously improve processes, tooling, and automation
- Collaborate closely with software development, security, customer support, and infrastructure teams
- Provide technical guidance and mentoring to operational staff
What You’re Great At
- Leading by example in a hands-on cloud operations leadership role
- Being self-reliant, proactive, and driven by technical excellence
- Communicating clearly and effectively with stakeholders across different geographic regions
- Providing technical directions to ensure stable, secure, and highly available operations
- Learning new technologies quickly and becoming a subject-matter expert when required
- Designing and maintaining effective monitoring and alerting for cloud services
What It Takes
- 5+ years’ experience in application management, systems administration, or cloud operations
- Tertiary qualification in Computer Science or a related discipline (preferred)
- Proven experience managing public cloud environments, with strong expertise in AWS
- Familiarity with the Fortify on Demand stack (highly desirable)
- Solid understanding of software architecture and DevOps principles
- Experience with infrastructure automation and scripting (e.g. PowerShell, Octopus Deploy, Terraform)
- Hands-on experience with monitoring and observability tools (e.g. AWS CloudWatch, Nagios, Grafana)
- Strong understanding of cloud security principles, backup strategies, and disaster recovery
- Working knowledge of networking concepts and protocols
Similar roles
-
Senior Site Reliability Engineer
Workday Sydney, New South Wales, Australia
-
Azure Site Reliability Engineer (SRE) - SaaS Operations
Zensar Bangalore South, Karnataka, India
-
Staff Software Engineer, Site Reliability Engineering Connections
Google Zurich, Zurich, Switzerland
-
Site Reliability Engineer
Xurrent Bangalore, Karnataka, India
-
Site Reliability Engineer
Cherokee Federal Boulder, Colorado, United States · $110K–$120K/yr
-
Principal Site Reliability Engineer (we have office locations in Cambridge, Leeds and London)
Genomics England London, England, United Kingdom · £107K/yr