Weekday AI

Lead SRE

Weekday AI Mumbai, Maharashtra, India

Technology, Information and Internet · 11-50 employees

8 h ago
sre Senior (5-10 yrs) Full-time India
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

You will lead the reliability strategy by designing multi-region AWS architectures and building self-healing systems. Additionally, you will mentor the SRE team, manage incident responses, and automate operational tasks to eliminate toil.

What they look for

SRE Monitoring CI/CD Infrastructure as Code Python Kubernetes AWS Terraform Bash ArgoCD Prometheus Grafana DevOps Platform Engineering Cloud Architecture GitOps

Requirements

Candidates must have at least 6 years of experience in DevOps or Platform Engineering with deep expertise in AWS and Kubernetes. Proficiency in IaC tools like Terraform and scripting languages such as Python or Bash is required.

Full description

This role is for one of Weekday’s clients Salary range: Rs 2000000 - Rs 2500000 (ie INR 20 - 25 LPA)

Min Experience: 6+ years Location: Mumbai, Maharashtra, India JobType: full-time

We are seeking a highly motivated and strategic-minded Cloud Engineer to join our dynamic team. You will lead our reliability strategy, build self-healing architecture, eliminate operational toil through software engineering, and mentor a high-performing team of SRE and Devops engineers.

Responsibilities:

Architecture & Reliability Ownership: Design multi-AZ and multi-region AWS architectures for zero-downtime releases. Establish and enforce SLIs, SLOs, and error budgets alongside product managers.

Infrastructure as Code (IaC): Standardize scalable cloud infrastructure using Terraform or AWS CDK. Ensure state isolation, modularity, and automated deployment pipelines.

Observability & Monitoring: Drive signal-to-noise alerting and build observability pipelines using Prometheus/Grafana, APM, AWS CloudWatch, or OpenTelemetry.

Incident Leadership & Postmortems: Serve as Incident Commander for high-severity issues. Lead blameless post-mortems and convert system failures into concrete engineering work.

Automation & Platform Tooling: Treat operations as a software problem by writing custom tooling in Bash or Python to eliminate manual operational toil.

Developer Experience & CI/CD: Maintain robust CI/CD pipelines (GitHub Actions, Tekton ArgoCD) to empower product developers with safe self-service deployments.

FinOps & Security: Partner with finance and security to optimize AWS spending (Savings Plans, right-sizing) while ensuring IAM policies, KMS encryption, and VPC configurations meet compliance standards.

Mentorship & On-Call Management: Set sustainable on-call rotation policies to prevent burnout and mentor junior and mid-level SREs.

 

Qualifications:

● Experience: 6+ years in DevOps, Platform Engineering, or Infrastructure.

● AWS Mastery: Deep hands-on experience across core AWS services (EKS, Lambda, RDS S3, Route 53, Transit Gateway).

● Kubernetes Expertise: Experience managing production EKS clusters, including workload scaling, ingress control, service meshes, and GitOps deployments (ArgoCD/Flux).

● Software Development: Strong coding skills in Bash or Python to build CLI tools, custom controllers, and API integrations.

● IaC Proficiency: Proven production usage of Terraform or AWS CDK.

● Observability: Track record of building metrics, logs, and trace pipelines from the ground up using modern observability stacks.

● Relevant certifications (e.g.,AWS Certified DevOps Engineer, Certified Kubernetes Administrator) are a plus.

Must-have skills

SRE, monitoring, CI/CD

Good-to-have skills

Infrastructure as Code, Python, Kubernetes

Similar roles