Senior Site Reliability Engineer
Sysco · Sri Lanka
Food and Beverage Services · 10,001+ employees
About the role
Design, build, and improve the reliability, scalability, and performance of enterprise cloud data platforms across GCP and AWS. Lead production incident response, root cause analysis, and drive cloud governance and FinOps initiatives.
What they look for
Requirements
Requires a bachelor's degree in Computer Science, Engineering, or an equivalent qualification. Candidates must have 3+ years of experience in SRE, DevOps, or Cloud Engineering with hands-on experience in cloud-native environments.
Benefits
Full description
JOB DESCRIPTION
Senior Site Reliability Engineer
The Big Picture
Sysco LABS is the Global In-House Center of Sysco Corporation (NYSE: SYY), the world’s largest foodservice company. Sysco ranks 56th in the Fortune 500 list and is the global leader in the trillion-dollar foodservice industry.
Sysco employs over 75,000 associates, has 337 smart distribution facilities worldwide and over 14,000 IoT-enabled trucks serving 730,000 customer locations. For fiscal year 2025 that ended June 29, 2025, the company generated sales of more than $81.4 billion.
Sysco LABS Sri Lanka delivers the technology that powers Sysco’s end-to-end operations.
Sysco LABS’ enterprise technology is present in the end-to-end foodservice journey, enabling the sourcing of food products, merchandising, storage and warehouse operations, order placement and pricing algorithms, the delivery of food and supplies to Sysco’s global network and the in-restaurant dining experience of the end-customer.
The Opportunity
We are currently on the lookout for a Senior Site Reliability Engineer to join our team.
Responsibilities:
- Design, build, and continuously improve the reliability, availability, scalability, and performance of enterprise cloud data platforms across Google Cloud Platform (GCP) and Amazon Web Services (AWS).
- Deploy, automate, and manage cloud infrastructure using Infrastructure as Code (Terraform preferred) and modern DevOps practices.
- Build and enhance observability across cloud infrastructure, databases, and data pipelines using Datadog and cloud-native monitoring solutions.
- Define, monitor, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), Error Budgets, and operational health metrics.
- Partner with Data Engineering teams to improve data platform reliability, data quality, disaster recovery, and operational resilience through proactive monitoring and early issue detection.
- Develop automation and self-healing solutions to reduce operational toil, streamline repetitive tasks, and improve engineering efficiency.
- Lead production incident response, root cause analysis (RCA), and post-incident reviews, driving permanent improvements to platform reliability.
- Drive cloud governance and FinOps initiatives by optimizing resource utilization, cloud costs, and operational best practices across GCP and AWS.
- Evaluate and introduce modern SRE, DevOps, and cloud technologies that improve platform reliability, operational maturity, and engineering productivity.
- Create and maintain operational documentation, runbooks, and recovery procedures while mentoring engineers and promoting Site Reliability Engineering best practices.
Requirements:
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or an equivalent qualification.
- 3+ years of experience in Site Reliability Engineering, DevOps, Cloud Engineering, Platform Engineering, or a similar role supporting enterprise production environments.
- Hands-on experience with Google Cloud Platform (preferred) and/or Amazon Web Services.
- Experience supporting cloud-native data platforms and services such as BigQuery, Cloud Composer, Pub/Sub, Cloud Run, Cloud Functions, DataStream, and Cloud Storage.
- Experience with Infrastructure as Code using Terraform (preferred) or similar technologies.
- Experience with Linux Administration, CI/CD pipelines, container platforms (Kubernetes/Docker), and automation using Python, Bash, or similar scripting languages.
- Hands-on experience with observability and monitoring platforms such as Datadog, Cloud Monitoring, Prometheus, or Grafana.
- Strong understanding of cloud security, IAM, FinOps principles, and cloud cost optimization best practices, and data reliability principles.
- Proven experience in incident management, root cause analysis, and driving reliability improvements in production environments.
- Excellent communication, collaboration, and documentation skills, with the ability to mentor engineers and promote SRE best practices.
Benefits
- Performance-based annual bonus
- Performance rewards and recognition
- Agile Benefits - special allowances for Health, Wellness & Academic purposes
- Paid birthday leave Team engagement allowance
- Comprehensive Health & Life Insurance Cover - extendable to parents and in-laws
- Hybrid work arrangement
Sysco LABS is an Equal Opportunity Employer.