Site Reliability Engineer - SRE Intermediate
FinThrive Gurugram, Haryana, India
Hospitals and Health Care · 1,001-5,000 employees
About the role
The Site Reliability Engineer will design, operate, and optimize cloud-native platforms on Azure while ensuring high availability and system resilience. They will lead incident response, perform root cause analysis, and implement automation to reduce operational toil.
What they look for
Requirements
Candidates must have 3–5+ years of experience in SRE or cloud engineering with strong expertise in Azure environments and Infrastructure as Code. A bachelor's degree in Computer Science or a related field is required, along with proficiency in monitoring tools and CI/CD pipelines.
Benefits
Full description
Site Reliability Engineer (SRE) – Cloud & Platform Engineering
Professional Summary
Site Reliability Engineer with 3–5+ years of experience in designing, operating, and optimizing cloud-native platforms with a strong focus on Azure environments. Proven expertise in building highly available, scalable, and secure systems using Infrastructure as Code (IaC) and automation-first practices.
Experienced in managing application hosting architectures including Azure App Services, ASEv3, Application Gateway (AGW) and Azure Front Door, ensuring high performance and resilience across distributed systems.
Demonstrates an automation mindset by leveraging modern engineering tools and AI-assisted development platforms (e.g., GitHub Copilot, Microsoft Copilot) to accelerate delivery, reduce operational toil, and improve reliability standards — with careful validation of outputs for security and production readiness.
Core Competencies
Cloud & Platform Engineering
• Microsoft Azure (Preferred)
• Understanding and experience in developing Azure function Apps, Azure logic Apps
• Understanding of event triggers, event hub, service bus.
• Application Hosting: App Services, App Service Plans, ASEv3
• Networking: Azure Application Gateway (AGW), Azure Front Door, VNet, NSGs, Load Balancing
• Cloud Architecture: High Availability, Fault Tolerance, Scalability Patterns
Incident Management and RCA
• Incident Management, P1 troubleshooting, Change Management
• Experienced in leading RCA and representing on the weekly call
• SLA / SLO / Error Budget concepts
• System Performance Optimization & Capacity Planning
• Toil Reduction through Automation
Infrastructure as Code & Automation
• Terraform, Azure Bicep, ARM Templates
• Azure Automation (Hybrid Workers)
• Azure Functions (Serverless automation)
• API-based automation and orchestration
Observability & Monitoring
• Azure Monitor, Log Analytics Workspace, Grafana, Site 24x7 (or similar SaaS based synthetic monitoring tool)
• Application Insights
• KQL (Kusto Query Language)
• Alert tuning and signal-to-noise optimization
AI-Enabled Productivity (Not as Skill)
• Leveraging GitHub Copilot / Microsoft Copilot for:
• Code acceleration and script generation
• Automation development support
• Troubleshooting and log analysis assistance
• Proven track record of workforce optimization leveraging AI tools.
• Applying validation frameworks to ensure secure, accurate, and production-grade outputs
DevOps & Integration
• CI/CD using Azure DevOps
• Deep understanding on version control
• API integrations (REST, Postman, SoapUI)
• Source control and release management
Professional Experience
Site Reliability Engineer / Cloud Engineer
SRE & Reliability Engineering
• Managed production environments ensuring high availability and reliability of cloud-hosted applications
• Led incident response, performed deep root cause analysis, and implemented preventive measures to reduce recurrence
• Improved system resilience through proactive monitoring and performance tuning strategies
Azure Application & Platform Engineering
• Designed and supported application architectures using:
• Azure App Services and App Service Plans
• Azure App Service Environment v3 (ASEv3) for isolated, high-scale workloads
• Azure Application Gateway (WAF-enabled) for L7 traffic management
• Azure Front Door for global traffic routing and failover
• Implemented secure and scalable cloud networking patterns, optimizing latency and throughput
Automation & Toil Reduction
• Identified repetitive operational tasks and reduced manual effort through automation-first solutions
• Developed automation using:
• Terraform / Bicep / ARM templates
• Azure Automation (Hybrid Workers)
• Azure Functions for event-driven workflows
• Leveraged AI-assisted tools (GitHub Copilot, Copilot) to accelerate scripting and automation development, while ensuring strict validation for enterprise use
Observability & Monitoring
• Built and enhanced observability using:
• Azure Monitor, Application Insights, Log Analytics
• Created KQL-based queries and dashboards for proactive issue detection
• Reduced false alerts by optimizing alert thresholds and improving signal quality
Performance & System Optimization
• Analyzed application performance across distributed systems to identify bottlenecks
• Implemented improvements through:
• Scaling strategies (horizontal & vertical)
• Network optimization (AGW / Front Door tuning)
• Backend service improvements
Collaboration & Engineering Enablement
• Partnered with SRE, CloudOps, and development teams to design resilient systems
• Contributed to runbooks, documentation, and operational standards
• Enabled engineering teams by improving platform reliability and deployment pipelines
Key Achievements
• Reduced manual operational effort by X% through automation initiatives
• Improved system availability to 99.X% by strengthening monitoring and failure handling mechanisms
• Decreased incident resolution time by X% via enhanced observability and streamlined runbooks
• Optimized application performance using Front Door and AGW tuning, reducing latency by X%
Education
• Bachelor’s Degree in Computer Science / Engineering or related field
Preferred/Additional Experience
• Experience with microservices and distributed architectures
• Exposure to low-code automation platforms
• Working knowledge of AWS cloud services
Preferred/Additional Certifications
• AZ-104 Azure Administrator
• AZ-700 Designing and Implementing Microsoft Azure Networking Solutions
• AZ-400 Microsoft Certified: DevOps Engineer Expert
About FinThrive FinThrive is advancing the healthcare economy. For the most recent information on FinThrive’s vision for healthcare revenue management visit finthrive.com/why-finthrive
Award-winning Culture of Customer-centricity and Reliability At FinThrive we’re proud of our agile and committed culture, which makes FinThrive an exceptional place to work. Explore our latest workplace recognitions at https://finthrive.com/careers#culture
Our Perks and Benefits FinThrive is committed to continually enhancing the colleague experience by actively seeking new perks and benefits.
· Professional development opportunities
· Term life, Accidental & Medical Insurance
· Meal and Transport arrangements
FinThrive’s Core Values and Expectations
· Demonstrate integrity and ethics in day-to-day tasks and decision-making, adhere to FinThrive’s core values of being Customer-Centric, Agile, Reliable, and Engaged, operate effectively in the FinThrive environment and the environment of the workgroup, maintain a focus on self-development and seek out continuous feedback and learning opportunities
· Support FinThrive’s Compliance Program by adhering to policies and procedures about HIPAA, GLBA, FCRA, and other laws applicable to FinThrive’s business practices; this includes becoming familiar with FinThrive’s Code of Ethics, attending training as required, notifying management or FinThrive’s Helpline when there is a compliance concern or incident, HIPAA-compliant handling of patient information, and demonstrable awareness of confidentiality obligations.
FinThrive is an Equal Opportunity Employer and ensures its employment decisions comply with principles embodied in Title VII, the Age Discrimination in Employment Act, the Rehabilitation Act of 1973, the Vietnam Veterans Readjustment Assistance Act of 1974, Executive Order 11246, Revised Order Number 4, and applicable state regulations.
© 2024 FinThrive. All rights reserved. The FinThrive name, products, associated trademarks, and logos are owned by FinThrive or related entities. RV092724TJO
Similar roles
-
Director, Site Reliability Engineering
Anduril Industries Costa Mesa, California, United States · $253K–$336K/yr
-
Senior Site Reliability Engineer (SRE)
Mirantis Sofia, Sofia-City, Bulgaria
-
Software Engineer III, Site Reliability Engineering
Google Pittsburgh, Pennsylvania, United States · $147K–$210K/yr
-
Senior Software Engineer, Site Reliability Engineering
Google Kirkland, Washington, United States · $174K–$252K/yr
-
Cloud Engineer / Site Reliability Engineer (SRE)
DFDS Denmark Copenhagen, Capital Region of Denmark, Denmark
-
Senior Software Developer, Site Reliability
Google Waterloo, Ontario, Canada · CA$182K–CA$186K/yr