Lead Site Reliability Engineer, Incident Management
Qualys pune, Maharashtra, India
Computer and Network Security · 1,001-5,000 employees
About the role
The Lead Site Reliability Engineer will provide technical leadership for production reliability, major incident management, and automation across the global SaaS platform. They will drive cross-functional initiatives, mentor engineering teams, and establish operational standards to improve system scalability and customer experience.
What they look for
Requirements
Candidates must have a bachelor's degree and 8–12+ years of experience in SRE or production operations. Proficiency in Linux, Kubernetes, cloud infrastructure, and programming languages like Python or Go is required, along with strong leadership and incident management skills.
Full description
Come work at a place where innovation and teamwork come together to support the most exciting missions in the world!
THE OPPORTUNITY
The Lead Site Reliability Engineer – Incident Management provides technical leadership for production reliability, major incident management, automation, and operational excellence across Qualys' global SaaS platform. This role owns complex cross-functional initiatives, drives engineering best practices, mentors engineers, and partners with architecture and product teams to improve reliability, scalability, and customer experience.
WHERE THIS ROLE SITS
Partners with Engineering, Infrastructure, DevOps, Cloud Operations, Database Engineering, Security, Product Engineering, Customer Support, and Executive Leadership.
WHAT YOU WILL DO
Reliability & Incident Management
· Lead enterprise-wide critical incident response and serve as Incident Commander.
· Own end-to-end restoration strategy for complex production outages.
· Drive executive communications and customer-impact assessments.
· Lead root cause analysis and systemic reliability improvements.
Reliability Engineering & Automation
· Design self-healing platforms and automation.
· Improve observability, SLOs, SLIs, and error budgets.
· Reduce MTTD and MTTR through engineering improvements.
Operational Excellence
· Lead capacity planning and operational readiness.
· Review architecture for reliability and scalability.
· Drive cross-functional operational standards.
Leadership & Collaboration
· Mentor Senior and Lead SREs.
· Drive SRE strategy and reliability roadmap.
· Influence engineering priorities and reliability culture.
WHAT GOOD LOOKS LIKE
· Drives measurable improvements in availability and resiliency.
· Builds a culture of automation and operational excellence.
· Leads major incidents with confidence.
· Influences engineering decisions across organizations.
DISTINGUISHING EXPECTATION
· Enterprise-wide reliability improvements
· Reduction in critical incidents
· Higher automation adoption
· Leadership across SRE organization
REQUIRED QUALIFICATIONS
· Bachelor's degree or equivalent.
· 8–12+ years in SRE/Production Operations.
· Strong Linux, Kubernetes, Cloud, Networking and Databases.
· Expertise in Python, Go, Bash or Java.
· Experience leading major incidents and mentoring engineers.
· Excellent communication and leadership skills.
PREFERRED QUALIFICATIONS
· Large-scale SaaS experience
· Multi-cloud expertise
· Chaos Engineering
· Cloud/Kubernetes certifications
· ITIL certification
WORK ENVIRONMENT
· Full-time
· 24x7 production support
· On-call leadership
· Hybrid/Remote
· Cross-functional collaboration
Similar roles
-
Lead SRE
66degrees Bengaluru, Karnataka, India
-
Senior Engineering Program Manager - Global SRE
Apple London, England, United Kingdom
-
Site Reliability Engineer (m/w/d) – Konstanz, Berlin oder remote
Karriere - SEITENBAU Konstanz, Baden-Württemberg, Germany
-
Principal Engineer - SRE
Arcesium LLC Bengaluru, Karnataka, India
-
Staff Site Reliability Engineer
Okta Bengaluru, Karnataka, India
-
Senior SRE
CloudRaft India