Principal SRE
Azira Bengaluru, Karnataka, India
Technology, Information and Media · 201-500 employees
About the role
Define and evolve the cloud infrastructure and reliability strategy to support global products, AI initiatives, and data platforms. Lead technical responses to complex production incidents while mentoring engineers and improving infrastructure standards across the organization.
What they look for
Requirements
Requires 12-15 years of experience in Site Reliability Engineering or cloud infrastructure at a Principal or Architect level. Candidates must possess deep expertise in AWS, Kubernetes, and Infrastructure-as-Code tools like Terraform, along with strong programming skills in Python or Go.
Full description
Why This Role Matters
This role will contribute to the Development and Improvement of Azira’s Cloud infrastructure, addressing key areas such as scalability, observability, security, and cost efficiency. The role will collaborate with teams across Engineering, AI, Product, Data, and Security to improve infrastructure standards, solve complex reliability challenges, and provide technical guidance and mentorship across the organization.
What you’ll do
- Define and evolve Azira’s cloud infrastructure and reliability strategy, ensuring it supports global products, data platforms, and AI initiatives.
- Design and implement scalable, resilient, secure, and highly available systems, including architecture standards, best practices, and disaster-recovery capabilities.
- Improve the reliability and performance of distributed, high-volume production systems by identifying and addressing single points of failure, capacity constraints, and recurring sources of instability.
- Develop and mature service-level indicators (SLIs), service-level objectives (SLOs), and error-budget practices to establish measurable reliability standards.
- Strengthen observability across applications and infrastructure through effective monitoring, logging, tracing, and alerting.
- Lead the technical response to complex production incidents, contribute as a senior escalation point when needed, and ensure incident learnings result in lasting improvements.
- Advance infrastructure-as-code, automation, and environment management practices to improve consistency, repeatability, scalability, and operational efficiency.
- Partner with Engineering, Data, AI, Product, and Security teams to embed reliability, security, and operational readiness throughout the development lifecycle, including supporting the infrastructure needs of AI workloads.
- Improve cloud cost visibility and efficiency while balancing performance, capacity, scalability, and spend; strengthen business continuity, backup, and disaster-recovery practices.
- Mentor SREs and engineers, raise technical standards, and communicate infrastructure risks, trade-offs, and recommendations clearly to technical leaders and business stakeholders.
What you Bring
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field.
- 12–15 years of experience in Site Reliability Engineering, Cloud Infrastructure, Platform Architecture, or related roles, with experience operating at a Principal, Staff-plus, Architect, or equivalent level.
- Significant experience in Site Reliability Engineering, cloud infrastructure, platform engineering, DevOps, or a related discipline, with experience operating at a Principal, Staff-plus, Architect, or equivalent level.
- Deep expertise in AWS, including experience architecting and operating infrastructure for products involving big data, high-volume workloads, or distributed systems. Experience with GCP in similar environments is also valuable.
- Strong understanding of cloud-native architecture, including Kubernetes, containers, networking, Linux, storage, databases, Amazon EMR, and cloud security.
- Advanced experience with Infrastructure-as-Code tools such as Terraform or OpenTofu, with a focus on building consistent, repeatable, and scalable infrastructure.
- Strong scripting and programming skills in Python, Bash, Go, or comparable languages, along with experience working with CI/CD and source-control platforms such as GitHub or Bitbucket.
- Experience with centralized authentication and authorization systems, including concepts such as OIDC and RBAC, as well as a solid understanding of infrastructure-level security and compliance requirements.
- Experience designing and operating distributed, data-intensive, or highly available systems at scale, with a strong understanding of observability across metrics, logs, traces, and alerting.
- Proven experience leading complex incident response, root-cause analysis, and reliability improvement initiatives, with the ability to balance immediate operational needs with long-term architectural improvements.
- Strong technical judgment and communication skills, with the ability to evaluate trade-offs across reliability, performance, security, speed, and cost, and communicate recommendations clearly across teams and regions.
- A collaborative leadership approach with a track record of mentoring engineers, influencing without formal authority, constructively challenging existing approaches, and taking ownership of problems through to resolution.
Similar roles
-
Site Reliability Engineer (m/f/x)
marbis GmbH Karlsruhe, Baden-Württemberg, Germany
-
Senior SRE
Banyan Software Canada · $145K–$170K/yr
-
Senior Staff Engineer, SRE
Aiven Helsinki, Uusimaa, Finland
-
Technical Program Manager, Networking SRE
Google Dublin, Leinster, Ireland · €94K–€96K/yr
-
Site Reliability Engineer III
American Express Bengaluru, Karnataka, India
-
Manager, Site Reliability Engineering
SolarWinds Brno, Southeast, Czechia