iScale Solutions

Lead Site Reliability Engineer

iScale Solutions

IT Services and IT Consulting · 501-1,000 employees

2 d ago
Remote sre Senior (5-10 yrs) Full-time
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

The Lead Site Reliability Engineer will manage service reliability through SLOs, SLIs, and error budgets while leading incident response and post-incident reviews. They will also design scalable infrastructure using IaC and implement comprehensive observability across cloud-based systems.

What they look for

Site Reliability Engineering Linux Networking AWS Kubernetes Python Go Bash OpenTelemetry Terraform Pulumi Chaos Engineering Observability Infrastructure as Code Incident Response Capacity Planning

Requirements

Candidates must have 5+ years of experience in SRE roles with deep knowledge of Linux, cloud platforms like AWS, and Kubernetes. Proficiency in Python or Go, shell scripting, and experience with observability and chaos engineering tools are required.

Benefits

Competitive Salary Package Vacation Leave Sick Leave Medical Insurance Dental Insurance Vision Insurance Statutory Benefits Training Certifications Mentorship Team-building Activities Wellness Programs Career Growth Opportunities Referral Rewards

Full description

This is a remote position.

• Implement and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets to drive reliability efforts.

• Develop systems that are resilient to failures and ensure 99.9%+ uptime for critical services.

• Lead incident response and post-incident reviews (blameless postmortems), ensuring robust root cause analysis and continuous improvement of systems.

• Automate incident detection and response using automated runbooks or predefined workflows.

• Write software as needed to support reliability or efficiency needs.

• Design and implement full observability across systems using modern tools like Open Telemetry for tracing, metrics, and logging.

• Use capacity planning, forecasting, and performance testing to ensure that the systems scale effectively as the user base and load grow.

• Collaborate with development and operations teams on building reliable, scalable, and high-performance services.

• Ensure best practices are followed across infrastructure design, deployment, and maintenance using tools like AWS, Kubernetes, EKS, Fargate, etc.

• Champion Infrastructure as Code (IaC) to provision, manage, and scale infrastructure using tools like Terraform, Pulumi, or similar.

• Get involved in chaos engineering initiatives.

• Participate on our on-call rotation.

• Drive advanced alerting and anomaly detection applied to metrics

Requirements

• 3-5 years of experience as an SRE, working in cloud-based environments and on-prem environments.

• Deep understanding of Linux systems, networking, and systems administration.

• Experience with cloud platforms like AWS, with a strong understanding of Kubernetes and container orchestration tools.

• Hands-on experience with observability tools such as Honeycomb, Grafana, Prometheus, Thanos, ELK (Elastic Stack), or Loki.

• Strong skills in at least one programming language (Python, Go) to write production level code.

• Strong skills in shell scripting using bash or similar. • Experience with OpenTelemetry or other distributed tracing systems, including tracing, metrics, and logs integration.

• Experience with Chaos Engineering methodologies and tools (Chaos Mesh, chaos monkey, AWS Fault Injection Simulator, etc.

Benefits

• Competitive Salary Package: Receive a pay package that matches your skills and experience.

• Vacation and Sick Leave credits: Enjoy vacation and sick leave credits to maintain work-life balance.

• Health Coverage: Get medical, dental, and vision insurance for you and your dependents.

• Government-Mandated Benefits: Full coverage of all statutory benefits like SSS, PhilHealth, and Pag-IBIG.

• Learning Opportunities: Access training, certifications, and mentorship to grow your career.

• Team Engagement: Join team-building activities and wellness programs.

• Modern Tools: Use the latest technology to excel in your role.

• Career Growth: Clear paths for promotion and professional development.

• Inclusive Culture: Be part of a diverse, supportive, and collaborative global team.

• Referral Rewards: Earn bonuses for bringing great talent to the team.

Similar roles