Manager, Site Reliability Engineering
Qore Technologies Limited Vientiane Capital, Vientiane Prefecture, Laos
Technology, Information and Internet · 11-50 employees
Applying here? Try the free cover letter tool — paste this posting and your résumé, no account needed.
About the role
The Manager of Site Reliability Engineering will lead and mentor the SRE team while defining strategy, roadmap, and best practices for system reliability. They are responsible for maintaining 99.9%+ availability, managing incident response, and overseeing infrastructure automation and security compliance.
What they look for
Requirements
Candidates must have at least 5 years of experience in DevOps, SRE, or Infrastructure Engineering, including 2 years in a leadership role. Proficiency in cloud platforms like AWS or Azure, container orchestration, and scripting languages is required, along with experience in high-transaction environments.
Benefits
Full description
1. Leadership & Strategy
- Lead,
mentor, and grow the SRE team; set clear goals, on-call structure, and career paths.
- Define
and own SRE strategy, roadmap, and best practices aligned with business and compliance requirements.
- Drive
a culture of reliability, automation, and blameless postmortems.
2. Reliability & Availability
- Own
SLAs, SLOs, and SLIs for all production platforms (core banking, APIs, payments).
- Ensure
99.9%+ availability of critical services and lead efforts to eliminate single points of failure.
- Manage
capacity planning, scalability, and disaster recovery (DR/BCP) strategies.
3. Infrastructure & Automation
- Own
and evolve our cloud and on-prem infrastructure (AWS/Azure, Kubernetes, Docker, Terraform).
- Drive
Infrastructure as Code (IaC), CI/CD, and GitOps maturity to enable safe, frequent releases.
- Lead
automation of operational toil, provisioning, and configuration management.
4. Incident & Problem Management
- Own
the incident response lifecycle - detection, escalation, resolution, and post-incident review.
- Build
and improve monitoring, alerting, logging, and observability stacks (Prometheus, Grafana, ELK/Datadog, PagerDuty).
- Act
as final escalation for P1/P2 incidents.
5. Security & Compliance
- Partner
with Security and Compliance to ensure infrastructure meets PCI-DSS, NDPA, CBN, and ISO 27001 requirements.
- Embed
security, secrets management, and vulnerability remediation into SRE practices.
- Own
change management and audit readiness for infrastructure changes.
6. Collaboration
- Collaborate
closely with Software Engineering, Product, Security, and Client Success to ensure reliability is built-in.
- Provide
technical guidance to engineering teams on resilient architecture patterns.
Requirements
Experience
- 5+
years in DevOps / SRE / Infrastructure Engineering, with 2+ years in a team lead role.
- Proven
experience managing highly available, high-transaction systems in fintech, banking, or large-scale B2B SaaS.
- Strong
track record managing production incidents and on-call teams.
Technical Skills
- Deep
expertise in Linux, networking, and distributed systems.
- Strong
hands-on experience with AWS (EC2, EKS, RDS, VPC, IAM, CloudWatch) or Azure.
- Expert
in Kubernetes, Docker, Terraform, or Ansible.
- Proficiency
in at least one scripting/programming language: Python, Go, or Bash.
- Experience
with CI/CD tools (Jenkins, GitLab CI, GitHub Actions, ArgoCD).
- Solid
understanding of monitoring/observability tools (Prometheus, Grafana, ELK, Datadog, New Relic).
Nice to Have
- Experience
with core banking systems, payment switches, or ISO 8583.
- Experience
with database reliability (PostgreSQL, MySQL, MongoDB, Redis).
- Certifications:
AWS Solutions Architect / DevOps Engineer, CKA/CKAD.
- Experience
with service mesh (Istio/Linkerd) and chaos engineering.
Soft Skills
- Excellent
leadership, communication, and stakeholder management.
- Strong
analytical and problem-solving mindset under pressure.
- Ability
to balance operational rigor with delivery speed.
Benefits
Qore provides the rare opportunity to make history in the financial space for Africa by Africans, while working with the smartest, brightest & coolest minds in Africa. Our people & culture team continuously thinks of innovative ways to improve employee experience and some of the other benefits of working with Qore includes:
- Very competitive and rewarding pay
- Flexible work option (i.e., Remote work)
- Paid Lunch for onsite work
- Lifelong Learnings
Similar roles
-
Software Developer III, Site Reliability
Google Waterloo, Ontario, Canada · CA$150K–CA$153K/yr
-
Senior Software Engineer, Site Reliability Engineering
Google New York, New York, United States · $174K–$252K/yr
-
Senior Site Reliability Engineer, Platform & Reliability (M/F/X)
Crossbeam Paris, Ile-de-France, France
-
Site Reliability Engineer (SRE) – Building X
Siemens Pune, Maharashtra, India
-
Head of Infrastructure SRE in China
Apple Shanghai, Shanghai, China
-
Site Reliability Engineering Intern (Brno, Czech Republic)
Red Hat Brno, Southeast, Czechia