Lead Site Reliability Engineer
IT Services and IT Consulting · 501-1,000 employees
Applying here? Try the free cover letter tool — paste this posting and your résumé, no account needed.
About the role
The Lead Site Reliability Engineer will manage service reliability through SLOs, SLIs, and error budgets while leading incident response and post-incident reviews. They will also design scalable infrastructure using IaC and implement comprehensive observability across cloud-based systems.
What they look for
Requirements
Candidates must have 5+ years of experience in SRE roles with deep knowledge of Linux, cloud platforms like AWS, and Kubernetes. Proficiency in Python or Go, shell scripting, and experience with observability and chaos engineering tools are required.
Benefits
Full description
This is a remote position.
• Implement and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets to drive reliability efforts.
• Develop systems that are resilient to failures and ensure 99.9%+ uptime for critical services.
• Lead incident response and post-incident reviews (blameless postmortems), ensuring robust root cause analysis and continuous improvement of systems.
• Automate incident detection and response using automated runbooks or predefined workflows.
• Write software as needed to support reliability or efficiency needs.
• Design and implement full observability across systems using modern tools like Open Telemetry for tracing, metrics, and logging.
• Use capacity planning, forecasting, and performance testing to ensure that the systems scale effectively as the user base and load grow.
• Collaborate with development and operations teams on building reliable, scalable, and high-performance services.
• Ensure best practices are followed across infrastructure design, deployment, and maintenance using tools like AWS, Kubernetes, EKS, Fargate, etc.
• Champion Infrastructure as Code (IaC) to provision, manage, and scale infrastructure using tools like Terraform, Pulumi, or similar.
• Get involved in chaos engineering initiatives.
• Participate on our on-call rotation.
• Drive advanced alerting and anomaly detection applied to metrics
Requirements
• 3-5 years of experience as an SRE, working in cloud-based environments and on-prem environments.
• Deep understanding of Linux systems, networking, and systems administration.
• Experience with cloud platforms like AWS, with a strong understanding of Kubernetes and container orchestration tools.
• Hands-on experience with observability tools such as Honeycomb, Grafana, Prometheus, Thanos, ELK (Elastic Stack), or Loki.
• Strong skills in at least one programming language (Python, Go) to write production level code.
• Strong skills in shell scripting using bash or similar. • Experience with OpenTelemetry or other distributed tracing systems, including tracing, metrics, and logs integration.
• Experience with Chaos Engineering methodologies and tools (Chaos Mesh, chaos monkey, AWS Fault Injection Simulator, etc.
Benefits
• Competitive Salary Package: Receive a pay package that matches your skills and experience.
• Vacation and Sick Leave credits: Enjoy vacation and sick leave credits to maintain work-life balance.
• Health Coverage: Get medical, dental, and vision insurance for you and your dependents.
• Government-Mandated Benefits: Full coverage of all statutory benefits like SSS, PhilHealth, and Pag-IBIG.
• Learning Opportunities: Access training, certifications, and mentorship to grow your career.
• Team Engagement: Join team-building activities and wellness programs.
• Modern Tools: Use the latest technology to excel in your role.
• Career Growth: Clear paths for promotion and professional development.
• Inclusive Culture: Be part of a diverse, supportive, and collaborative global team.
• Referral Rewards: Earn bonuses for bringing great talent to the team.
Similar roles
-
Senior Site Reliability Engineer
Akamai Bengaluru, Karnataka, India
-
Site Reliability Engineer (SRE), London
Apple London, England, United Kingdom
-
Principal SRE Engineer
Entain Hyderabad, Telangana, India
-
IN_Manager_Site Reliability Engineering_GCC_Advisory_Bangalore
PwC Bengaluru, Karnataka, India
-
Site Reliability Engineer
CrelioHealth Pune, Maharashtra, India
-
Site Reliability Engineer I
CME Group Bengaluru, Karnataka, India