CBTS

Lead Engineer – Site Reliability Engineering

CBTS Chennai, Tamil Nadu, India

IT Services and IT Consulting · 1,001-5,000 employees

11 h ago
sre Senior (5-10 yrs) Full-time India
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

The Lead Engineer will ensure high availability, scalability, and resilience of cloud platforms by implementing SRE principles and automation. They will also lead incident response, drive observability improvements, and mentor engineering teams on operational excellence.

What they look for

Site Reliability Engineering Automation Observability Cloud Platforms Infrastructure Engineering Incident Response Performance Engineering Chaos Engineering Configuration Management Vulnerability Remediation Capacity Planning Fault Injection Runbooks SRE Frameworks Reliability Guardrails

Requirements

The role requires expertise in SRE frameworks, cloud infrastructure, and automation workflows. Candidates must be capable of defining reliability patterns, managing incident readiness, and ensuring security compliance across services.

Full description

Lead Engineer – Site Reliability Engineering Role Purpose : Ensures high availability, performance, scalability, and resilience of cloud and infrastructure platforms by applying SRE engineering principles, automation-first practices, observability, and continual reliability improvements across services and platforms.

Key Responsibilities:

·       Implement SRE frameworks, SLIs/SLOs/SLAs, error budgets, performance engineering, and reliability guardrails across cloud platforms & services.

·       Drive automation for provisioning, deployment, configuration management, drift control, patching, recovery, and operations workflows.

·       Build observability stack, dashboards, anomaly detection, synthetic tests, runbooks, incident readiness and RCA automation.

·       Partner with DevOps, Platform Engineering, Cloud Engineering & application squads to define reliability patterns, capacity planning & scalable workload landing models.

·       Lead incident response, major incident coordination, postmortem improvement actions, resiliency testing, fault injection, chaos engineering initiatives.

·       Ensure infra security alignment, vulnerability remediation, compliance & secure configuration baselines in cloud infrastructure.

·       Mentor engineers in SRE best practices, automation pipelines, tooling standardization and operational excellence.

Similar roles