Site Reliability Engineer
WorkSpan · Bengaluru, Karnataka, India
Software Development · 51-200 employees
About the role
Design, automate, and maintain resilient multi-cloud infrastructure across AWS, GCP, and Azure to ensure high availability and fault tolerance. Manage CI/CD pipelines, establish observability standards, and lead incident response to eliminate manual toil.
What they look for
Requirements
Requires 4+ years of experience in DevOps or SRE with hands-on production experience in at least two major cloud providers. Proficiency in Infrastructure as Code tools, containerization, and automation scripting is essential.
Full description
Job Brief: Multi-Cloud SRE / DevOps Engineer
We are seeking a versatile Multi-Cloud Site Reliability / DevOps Engineer to design, automate, and maintain our geographically distributed infrastructure across Amazon Web Services (AWS), Google Cloud Platform (GCP), and Microsoft Azure. In this role, you will champion a true multi-cloud strategy by avoiding vendor lock-in, optimizing cloud costs, and deploying highly available, fault-tolerant systems. You will act as the bridge between software development and IT operations, treating infrastructure as a software engineering problem to eliminate manual toil.
Core Responsibilities
- Multi-Cloud Architecture & Orchestration: Design, provision, and maintain resilient infrastructure across AWS, GCP, and Azure using cloud-agnostic Infrastructure as Code (IaC) tools to ensure environment parity.
- CI/CD Pipeline Management: Build, automate, and streamline cross-platform continuous integration and delivery pipelines that seamlessly route workloads to their optimal cloud environments.
- Site Reliability & Observability: Establish Service Level Indicators (SLIs) and Objectives (SLOs). Set up centralized, multi-cloud logging and monitoring platforms to trace metrics and proactively diagnose bottlenecks before they impact users.
- Incident Response & On-Call: Triage production issues, manage incident bridges, and lead blameless post-mortems to drive continuous systemic improvements.
- Disaster Recovery & High Availability: Engineer multi-region and active-active failover strategies between different public cloud providers to prevent single points of failure.
- FinOps & Resource Optimization: Monitor utilization patterns, implement automated scaling, and drive cost-saving initiatives across all three cloud platforms.
Cloud-Specific Engineering Duties
As a multi-cloud engineer, you will be expected to leverage and manage native services across the major providers:
Cloud Platform
Key Responsibilities & Target Services
AWS
Manage compute and orchestration using EKS, EC2, and Lambda. Configure network topology (VPC, Route53) and manage access via AWS IAM. Utilize native or agnostic tools (CloudFormation, Terraform) for provisioning.
GCP
Oversee containerized workloads via GKE (Google Kubernetes Engine). Optimize data pipelines/storage using BigQuery and Cloud Storage. Manage Google Cloud Load Balancing and Anthos for multi-cloud deployments.
Azure
Deploy and scale applications using AKS (Azure Kubernetes Service). Automate infrastructure management via ARM templates or Bicep. Ensure tight integration with Azure Active Directory (Entra ID) for secure identity management.
Required Qualifications & Skills
- Experience: 4+ years in DevOps, SRE, or Cloud Engineering, with hands-on production experience in at least two of the three major cloud providers (AWS, GCP, Azure).
- Infrastructure as Code (IaC): Advanced proficiency in Terraform (strongly preferred for multi-cloud parity), alongside familiarity with CloudFormation, ARM/Bicep, or Deployment Manager.
- Containerization & Kubernetes: Deep understanding of Docker and production-level experience administering Kubernetes clusters (EKS, GKE, AKS).
- Automation/Scripting: Strong programming skills in Python, Go, or Bash for automating systemic operations and interacting with various Cloud APIs.
- CI/CD Tooling: Experience with cloud-agnostic deployment tools like GitHub Actions, GitLab CI, ArgoCD, Jenkins, or CircleCI.
- Observability Stack: Hands-on experience configuring Prometheus, Grafana, Datadog, or the ELK/EFK stack for unified multi-cloud dashboards.
Preferred Qualifications
- Experience with Cloud Center of Excellence (CCoE) practices and multi-cloud governance frameworks.
- Certifications such as Certified Kubernetes Administrator (CKA), AWS Certified DevOps Engineer, Google Cloud Professional Cloud Architect, or Azure DevOps Engineer Expert.