About the role
You will provide reliability coverage and facilitate clean handoffs between US and India shifts for global restaurant technology. The role involves transforming reliability practices from reactive incident response toward a platform engineering model using self-service and auto-healing patterns.
What they look for
Requirements
Candidates must have 5+ years of experience in SRE, DevOps, or infrastructure roles with hands-on expertise in observability and automation. Proficiency in at least one scripting language and experience with cloud providers and container orchestration are required.
Full description
- 2+ years of experience in SRE, DevOps, production support, or infrastructure engineering roles
- Hands-on experience with monitoring and observability tooling (e.g., Datadog, Prometheus, Grafana, CloudWatch, or similar)
- Working knowledge of at least one major cloud provider (AWS preferred)
- Proficiency in at least one scripting or programming language (e.g., Python, Bash, Go) for automation, with demonstrated examples of automating away manual operational work
- Experience participating in incident response and on-call or shift-based operations
- Understanding of SLI/SLO concepts and reliability engineering fundamentals
- Ability to work follow-the-sun shift rotations, including structured handoffs with teams in other regions
- Strong written and verbal English communication skills for cross-region collaboration
Similar roles
-
Site Reliability Engineer
Akamai Bengaluru, Karnataka, India
-
Senior Site Reliability Engineer (Performance and Scalability)
Digital Zone Poland
-
Senior Site Reliability Engineer
TeamViewer Germany GmbH Austin, Texas, United States
-
Senior Site Reliability Engineer
Duplo Lagos, Lagos State, Nigeria
-
Sr Staff Site Reliability Engineer
Palo Alto Networks Sofia, Sofia-City, Bulgaria
-
Staff Site Reliability Engineer
CME Group Chicago, Illinois, United States · $132K–$220K/yr