Lead SRE & Support Engineer
Providence Hyderabad, Telangana, India
Hospitals and Health Care · 1,001-5,000 employees
About the role
The Lead SRE will maintain the reliability, availability, and operational performance of critical cloud-native services in Microsoft Azure. They will manage DevOps pipelines, drive incident resolution, and implement automation to improve service resilience and reduce operational toil.
What they look for
Requirements
Candidates must have a bachelor's degree and 6-9 years of experience in SRE, production support, or cloud operations. Proficiency in Azure, SQL, Snowflake, and scripting languages like Python or PowerShell is required, along with the ability to work a permanent night shift.
Full description
About the Team
The Cybersecurity Product Engineering Team designs, develops, and operates strategic platforms that support security posture visibility, operational intelligence, enterprise risk management, and secure engineering outcomes.
The team works at the intersection of cybersecurity, cloud engineering, data engineering, site reliability engineering, and artificial intelligence. Reliability and operational excellence are central to how we deliver secure, resilient, scalable services to global stakeholders.
About the Role
We are seeking a highly motivated Lead Site Reliability Engineer (SRE) Support Engineer (3P) to join the Cybersecurity Product (Software) Engineering Team. This role is responsible for maintaining the reliability, availability, and operational performance of critical cloud-native services running in Microsoft Azure currently, with potential to extend to AWS, GCP Cloud Service Providers.
This is a full-time night-shift role (9:00 PM to 6:00 AM IST) supporting US business hours. While office presence requirements can be discussed, candidates must be based in Hyderabad and available to work from the Hyderabad location as needed.
The successful candidate will partner with software engineers, platform engineers, data engineers, and cybersecurity stakeholders to support production services, manage incidents, improve observability, automate repetitive work, and drive continuous operational improvement. The ideal candidate combines strong troubleshooting skills with an ownership mindset and practical experience using AI-assisted engineering tools.
TECHNICAL JOB DESCRIPTION
- Own and manage DevOps pipelines for the platform across application code, data, RAG, and ML workloads; build and enhance pipelines as needed to improve reliability, automation, and deployment efficiency.
- Provide hands-on Production Support for product with clear, timely incident updates that summarize impact, investigation status, mitigation, and next actions.
- Drive incidents through containment, recovery, validation, closure, and handoff when cross-shift follow-up is required.
- Contribute to root cause analysis and track corrective and preventive actions to completion.
- Participate in a 24x7 support and on-call model as required by business and service needs.
Site Reliability Engineering
- Improve service reliability, resilience, scalability, and operational effectiveness through engineering-led support practices.
- Identify recurring failure patterns, reliability risks, and opportunities to reduce manual operational toil.
- Contribute to Service Level Indicators (SLIs), Service Level Objectives (SLOs), availability targets, and service health reviews.
- Partner with engineering teams on operational readiness, recovery procedures, capacity considerations, and platform hardening.
- Use incident and service-health learnings to recommend durable technical and process improvements.
Azure-Native Observability
- Use Azure Monitor, Application Insights, Log Analytics workspaces, Azure Alerts, and Azure Service Health to monitor and diagnose services.
- Investigate system behavior through metrics, logs, distributed traces, dependency maps, queries, and correlated events.
- Build and improve actionable dashboards, alert rules, health views, and operational reports.
- Tune monitoring and alerting to improve signal quality, reduce alert fatigue, and shorten detection and recovery cycles.
- Contribute to observability standards and consistent telemetry practices across supported services.
Data Platform Operations
- Support the operational reliability of cloud-based data services and analytics workloads.
- Troubleshoot Snowflake connectivity, access, workload, query performance, and operational issues within the scope of the support role.
- Investigate data ingestion, transformation, reporting, and pipeline failures in collaboration with data engineering teams.
- Validate data availability and operational recovery after incidents, releases, or maintenance activities.
- Use SQL to investigate data issues, validate processing outcomes, and support incident diagnosis.
Automation and Operational Excellence
- Develop and maintain Python, PowerShell, or shell-based automation for health checks, evidence gathering, diagnostics, and routine support activities.
- Create reusable utilities and workflow improvements that reduce manual effort and improve response consistency.
- Identify opportunities for safe self-service and self-healing capabilities with appropriate controls and auditability.
- Improve support processes through standardization, measurable outcomes, documentation, and continual learning.
- Contribute operational feedback to backlog prioritization and engineering improvement plans.
AI-Assisted Operations
- Use approved enterprise AI assistants and copilots to accelerate troubleshooting, knowledge retrieval, scripting, documentation, and incident summarization.
- Apply effective prompt engineering techniques to produce clear, context-aware operational outputs and refine results through validation.
- Use AI assistance to summarize logs and telemetry, detect patterns, propose hypotheses, and organize evidence while independently validating conclusions.
- Use AI-assisted coding tools to draft or improve scripts, tests, queries, and automation with appropriate review and secure coding practices.
- Create and improve runbooks, knowledge articles, incident timelines, and post-incident documentation using AI-assisted workflows.
- Understand foundational AIOps concepts such as anomaly detection, event correlation, alert enrichment, and intelligent triage.
- Protect confidential, personal, security-sensitive, and regulated information when using AI tools, following organizational data-handling requirements.
- Recognize AI limitations, including inaccurate or incomplete outputs, and apply human review before operational use.
Required Qualifications
- Bachelor's degree in Computer Science, Information Technology, Engineering, Cybersecurity, or a related technical discipline, or equivalent practical experience.
- 6-9 years of experience in Site Reliability Engineering, Production Support, Application Support, Platform Operations, or Cloud Operations.
- Hands-on experience supporting enterprise-scale production applications and services in Microsoft Azure. With experience/ exposure to other Cloud Service Providers (AWS, GCP).
- Experience with incident response, service restoration, root cause analysis, change support, and production release validation.
- Working knowledge of Azure-native monitoring and observability services.
- Hands-on SQL and Snowflake support or troubleshooting experience.
- Ability to automate operational work using Python, PowerShell, or shell scripting.
- Working knowledge of REST APIs, JSON, authentication and authorization concepts, and service connectivity troubleshooting.
- Experience with Azure DevOps, Git, and CI/CD operational support.
- Working knowledge of Linux and Windows environments.
- Strong analytical, problem-solving, written communication, and stakeholder coordination skills.
- Ability and willingness to work a permanent US shift and participate in a 24x7 or on-call support model as required.
Preferred Qualifications
- Experience supporting cybersecurity, risk management, governance, compliance, security analytics, or enterprise security platforms.
- Exposure to cloud-native, microservices-based, and distributed application architectures.
- Familiarity with containerized environments and Kubernetes or Azure Kubernetes Service concepts.
- Practical understanding of SRE principles, reliability metrics, SLI/SLO practices, and toil reduction.
- Experience working in Agile, DevOps, or product engineering delivery models.
- Experience using approved enterprise AI assistants such as Microsoft Copilot, GitHub Copilot, or comparable tools in engineering workflows.
- Familiarity with responsible AI usage, secure prompt practices, and validation of AI-generated technical output.
- Microsoft Azure or related cloud certifications.
Core Technical Capabilities
Capability
Expected Knowledge and Experience
Cloud Platform
Microsoft Azure (M), AWS (O), GCP (O) cloud-native application and platform support
Observability
Azure Monitor; Application Insights; Log Analytics; Azure Alerts; Azure Service Health
Reliability Operations
Incident and problem management; RCA; change support; release validation; SLI/SLO awareness
Data Technologies
Snowflake; SQL; data investigation; query and pipeline troubleshooting
Automation
Python; PowerShell; shell scripting; operational tooling and runbooks
Application Support
REST APIs; JSON; authentication and authorization; service connectivity
DevOps
Azure DevOps; Git; CI/CD operational support
Operating Systems
Linux; Windows Server fundamentals
AI for Operations
Copilots; prompt engineering; AI-assisted diagnostics, coding, documentation, and AIOps awareness
What Success Looks Like
• Maintain product and platform uptime by ensuring reliable operations, rapid incident resolution, resilient infrastructure, and highly automated deployment and monitoring processes.
- Independently manages production incidents and operational escalations during assigned support hours.
- Meets agreed response, restoration, communication, and resolution expectations while minimizing business impact.
CYBERSECURITY PRODUCT ENGINEERING | TECHNICAL JOB DESCRIPTION
- Improves monitoring coverage, alert quality, operational visibility, and service reliability.
- Reduces recurring incidents and manual effort through automation, preventive actions, and effective knowledge management.
- Produces clear incident records, actionable root cause analyses, and well-owned follow-up actions.
- Uses AI-assisted tools responsibly to improve productivity, troubleshooting quality, and documentation without compromising security or human oversight.
- Builds trusted working relationships across Cybersecurity, Product Engineering, Cloud, Platform, and Data teams.
Similar roles
-
Senior Site Reliability Engineer (Digital Infrastructure)
Egis Group Melbourne, Victoria, Australia
-
IT Site Reliability Engineer — API Management Platforms
Texas Instruments Dallas, Texas, United States
-
Staff Site Reliability Engineer, GovCloud
Medallia Mclean, Virginia, United States
-
Consultant Specialist(SRE)
HSBC Global Services Limited Tianhe District, Guangdong, China
-
Site Reliability Engineer
IFS Tokyo, Tokyo, Japan
-
Site Reliability Engineer
Finastra Mississauga, Ontario, Canada · CA$95K–CA$130K/yr