Lead SRE
Weekday AI Mumbai, Maharashtra, India
Technology, Information and Internet · 11-50 employees
About the role
You will lead the reliability strategy by designing multi-region AWS architectures and building self-healing systems. Additionally, you will mentor the SRE team, manage incident responses, and automate operational tasks to eliminate toil.
What they look for
Requirements
Candidates must have at least 6 years of experience in DevOps or Platform Engineering with deep expertise in AWS and Kubernetes. Proficiency in IaC tools like Terraform and scripting languages such as Python or Bash is required.
Full description
This role is for one of Weekday’s clients Salary range: Rs 2000000 - Rs 2500000 (ie INR 20 - 25 LPA)
Min Experience: 6+ years Location: Mumbai, Maharashtra, India JobType: full-time
We are seeking a highly motivated and strategic-minded Cloud Engineer to join our dynamic team. You will lead our reliability strategy, build self-healing architecture, eliminate operational toil through software engineering, and mentor a high-performing team of SRE and Devops engineers.
Responsibilities:
Architecture & Reliability Ownership: Design multi-AZ and multi-region AWS architectures for zero-downtime releases. Establish and enforce SLIs, SLOs, and error budgets alongside product managers.
Infrastructure as Code (IaC): Standardize scalable cloud infrastructure using Terraform or AWS CDK. Ensure state isolation, modularity, and automated deployment pipelines.
Observability & Monitoring: Drive signal-to-noise alerting and build observability pipelines using Prometheus/Grafana, APM, AWS CloudWatch, or OpenTelemetry.
Incident Leadership & Postmortems: Serve as Incident Commander for high-severity issues. Lead blameless post-mortems and convert system failures into concrete engineering work.
Automation & Platform Tooling: Treat operations as a software problem by writing custom tooling in Bash or Python to eliminate manual operational toil.
Developer Experience & CI/CD: Maintain robust CI/CD pipelines (GitHub Actions, Tekton ArgoCD) to empower product developers with safe self-service deployments.
FinOps & Security: Partner with finance and security to optimize AWS spending (Savings Plans, right-sizing) while ensuring IAM policies, KMS encryption, and VPC configurations meet compliance standards.
Mentorship & On-Call Management: Set sustainable on-call rotation policies to prevent burnout and mentor junior and mid-level SREs.
Qualifications:
● Experience: 6+ years in DevOps, Platform Engineering, or Infrastructure.
● AWS Mastery: Deep hands-on experience across core AWS services (EKS, Lambda, RDS S3, Route 53, Transit Gateway).
● Kubernetes Expertise: Experience managing production EKS clusters, including workload scaling, ingress control, service meshes, and GitOps deployments (ArgoCD/Flux).
● Software Development: Strong coding skills in Bash or Python to build CLI tools, custom controllers, and API integrations.
● IaC Proficiency: Proven production usage of Terraform or AWS CDK.
● Observability: Track record of building metrics, logs, and trace pipelines from the ground up using modern observability stacks.
● Relevant certifications (e.g.,AWS Certified DevOps Engineer, Certified Kubernetes Administrator) are a plus.
Must-have skills
SRE, monitoring, CI/CD
Good-to-have skills
Infrastructure as Code, Python, Kubernetes
Similar roles
-
Site Reliability Automation Engineer
Playtech Sofia, Sofia-City, Bulgaria
-
Software Engineer Lead - Site Reliability
Standard Life plc Telford, England, United Kingdom · £60K–£80K/yr
-
Engineering Manager - SRE
Plume Hyderabad, Telangana, India
-
Senior Site Reliability Engineer - Search
Algolia Paris, Ile-de-France, France · €70K–€97K/yr
-
Staff Systems Developer Manager, AlphaNet Core SRE
Google Waterloo, Ontario, Canada · CA$216K–CA$221K/yr
-
Site Reliability Engineer
Moniepoint South Africa