SYSTRA

Site Reliability Engineer (SRE) – Azure & DevOps

SYSTRA · Noida, Uttar Pradesh, India

Engineering Services · 10,001+ employees

20 h ago
Principal (10+ yrs) Full-time India
Log in to apply, save this posting, or score it against your profile with AI.

About the role

The Site Reliability Engineer will ensure the reliability, scalability, and performance of the Azure platform while acting as a technical bridge between development and infrastructure teams. Key duties include defining SLOs, managing incident responses, automating deployment workflows, and mentoring teams on cloud-native best practices.

What they look for

Azure DevOps Site Reliability Engineering Terraform Kubernetes Docker Azure DevOps GitLab CI Python Bash Observability FinOps DevSecOps Infrastructure as Code Incident Management Cloud Architecture

Requirements

The ideal candidate possesses 10 to 15 years of professional experience in high-stakes DevOps and SRE environments, with mandatory experience in a multinational company. A B.Tech or M.Tech degree in Computer Science or a related field is required, along with expert proficiency in Azure, Terraform, and container orchestration.

Full description

It has been more than 60 years since SYSTRA has garnered expertise that spans the entire spectrum of Mass Rapid Transit System. SYSTRA India’s valuable presence in India roots back to 1957, where SYSTRA worked on the electrification of Indian Railways. Our technical excellence, holistic approach and the tremendous talent provides a career that puts people who join us at the heart of improving transportation and urban infrastructure efficiency.

Understand better who we are by visiting www.systra.in

Context

The Site Reliability Engineer (SRE) – Azure & DevOps reports directly to the Group IT Hosting Manager and serves as the primary technical interface between Development teams, Hosting/Infrastructure teams, and DevOps engineers.

The main mission of this role is to ensure the reliability, scalability, performance, and operational excellence of the Azure platform supporting business-critical applications. The role combines strong Site Reliability Engineering practices with modern DevOps capabilities to improve platform resilience, delivery velocity, and technical maturity across the organization.

The position acts as a key technical referent for cloud-native practices and contributes to the continuous evolution of the platform by promoting automation, observability, security, and cost optimization. The engineer is also expected to elevate the capabilities of surrounding teams by mentoring developers, sharing best practices, and supporting the adoption of SRE, DevSecOps, and FinOps principles.

Role: Individual Contributor

Missions/Main Duties

Main Activities

  • Act as the main technical point of contact for Development teams on the Azure ecosystem and cloud platform best practices.
  • Bridge the gap between application teams and Hosting/Infrastructure teams by translating technical requirements in both directions.
  • Mentor and upskill developers on observability, resilience engineering, release reliability, and continuous deployment practices.
  • Promote and disseminate SRE, DevOps, DevSecOps, cloud-native architecture, and FinOps principles across engineering squads.
  • Guarantee the availability, stability, and performance of mission-critical production applications hosted on Azure.
  • Define, monitor, and communicate SLOs, SLIs, and Error Budgets in collaboration with product and technical teams.
  • Implement and improve observability solutions including logging, metrics, tracing, monitoring, and alerting.
  • Lead incident response activities including troubleshooting, service restoration, root cause analysis, and blameless post-mortems.
  • Develop and maintain Infrastructure as Code using Terraform for Azure environments.
  • Design, maintain, and optimize CI/CD pipelines using Azure DevOps, GitLab CI, or equivalent tooling.
  • Automate test environment provisioning and deployment workflows to improve consistency, speed, and reliability.
  • Industrialize deployment and operations practices to ensure reproducibility, traceability, compliance, and security.
  • Challenge and improve architectural choices together with development, architecture, and operations teams.
  • Actively contribute to FinOps initiatives by identifying and implementing cloud cost optimization opportunities.
  • Support DevSecOps initiatives by integrating security controls and best practices into delivery and operational processes.
  • Conduct technology watch and recommend innovative solutions to strengthen the Azure platform and engineering practices.
  • Produce and maintain high-quality technical documentation and share best practices across the engineering organization.
  • Collaborate with international teams and interact professionally with Microsoft Azure support when required.

Technical Skills

  • Azure Cloud
  • Advanced administration and operational knowledge of Azure core services
  • App Services
  • Container Apps
  • API Management (APIM)
  • Blob Storage
  • Cosmos DB
  • Azure Networking
  • Identity and Access Management (IAM)
  • Infrastructure as Code
  • Expert proficiency with Terraform
  • Reusable module design
  • Environment standardization
  • Infrastructure automation on Azure
  • Containerization & Orchestration
  • Strong hands-on experience with Docker
  • Production-grade Kubernetes administration
  • Azure Kubernetes Service (AKS)
  • Container deployment, scaling, and reliability practices
  • CI/CD & DevOps Tooling
  • Azure DevOps
  • GitLab CI or equivalent platforms
  • Pipeline design and optimization
  • Automated build, test, and deployment workflows
  • Scripting & Automation
  • Bash scripting
  • Python scripting
  • Automation of operational and deployment tasks
  • Observability & Monitoring
  • Elastic Stack (ELK)
  • Grafana
  • Prometheus
  • Centralized logging, metrics, tracing, alerting, and dashboards
  • SRE Practices
  • SLO / SLI definition and tracking
  • Error Budget management
  • Incident management and response
  • Root cause analysis
  • Toil reduction and operational excellence
  • Security & Optimization
  • DevSecOps practices
  • Cloud security integration
  • Cost governance and FinOps optimization
  • Reliability-focused architecture reviews

Profile/Skills

Profile & Personal Attributes

The ideal candidate holds a B.Tech or M.Tech degree in Computer Science, Engineering, or a related discipline. A Master’s degree or equivalent academic background is highly valued.

The candidate should have 10 to 15 years of professional experience, including substantial exposure to high-stakes DevOps and Site Reliability Engineering environments. Experience working in a multinational company (MNC) is mandatory, as the role requires collaboration with international teams and engagement with global support structures.

The successful profile combines strong technical depth with excellent interpersonal effectiveness. The candidate must be capable of communicating complex topics clearly to a wide range of stakeholders.

A proactive, self-driven, and solution-oriented attitude is essential. This person should be comfortable challenging existing practices, proposing improvements, and acting as a force multiplier across Dev, Ops, Architecture, and Product teams. Professional English proficiency is required for interactions with Microsoft Azure support and international colleagues.

This is an Individual Contributor role requiring strong ownership, influence, autonomy, and the ability to drive technical excellence through collaboration rather than hierarchical management.

We commit to put people who join us at the heart of improving transportation and urban infrastructure efficiency. As we are growing, this is time to be a part of this challenging adventure.It’s not a job - it’s a career!

Workplace Type

On-site