Senior Site Reliability Engineer
Technology, Information and Internet · 2-10 employees
About the role
You will design and evolve scalable GCP infrastructure while building internal tooling to improve developer productivity. Additionally, you will champion reliability practices, manage incident responses, and optimize cloud infrastructure costs.
What they look for
Requirements
Candidates must have at least 3 years of SRE or production experience with proficiency in GCP, Kubernetes, and Infrastructure as Code. Strong skills in observability stacks and scripting languages like Python, Bash, or Go are also required.
Full description
About the Role
We are a well-funded AI/ML company operating at the intersection of geospatial intelligence and climate technology. Our engineering team builds products that rely on robust, scalable cloud infrastructure — and we're looking for a Senior Site Reliability Engineer to help us evolve that foundation.
In this role, you'll own and advance our GCP infrastructure and DevOps practices, driving improvements in incident management, SLOs, error budgets, and observability. You'll partner closely with Product & Engineering teams to use DORA metrics as a lever for continuous improvement, while also helping the broader organization understand and optimize cloud costs.
What You'll Do
- Design and evolve our cloud infrastructure on GCP for scale and resilience.
- Build internal tooling and automation that promote team autonomy and developer productivity.
- Advance our observability platform — metrics, logging, tracing, and alerting — to reduce mean time to recovery.
- Build visibility into infrastructure costs and drive optimization initiatives.
- Champion reliability best practices across engineering, including SLOs/SLIs, error budgets, and post-incident reviews.
- Participate in on-call rotation and lead incident management efforts.
What We're Looking For
Required:
- 3+ years of Site Reliability Engineering or production SRE experience.
- Proficiency with Google Cloud Platform (GCP), including cost optimization and governance.
- Hands-on experience with Kubernetes for cluster and workload management.
- Infrastructure as Code experience using tools such as Terraform or Deployment Manager.
- Scripting and automation skills in Python, Bash, or Go.
- Strong observability stack experience: Prometheus, Grafana, OpenTelemetry, logging, and tracing.
- Demonstrated ability to define and implement SLOs, SLIs, and error budgets.
- Experience with incident management, post-incident reviews, and on-call rotation.
Nice to Have:
- Experience designing and evolving cloud infrastructure at scale.
- Familiarity with DORA metrics and using them to drive engineering effectiveness.
Please note: Visa sponsorship is not available for this role.
Location
This is a fully remote role open to candidates based in the EU, UK, or North America (including Canada, Denmark, Estonia, France, Netherlands, Portugal, Sweden, Switzerland, and the United Kingdom).
Compensation & Benefits
Compensation details were not specified for this role. We are happy to discuss salary expectations during the interview process.
Similar roles
-
Site Reliability Engineer - HM: Mukesh
NTT Data Singapore Singapore, Singapore
-
Site Reliability Engineer III - Machine Learning
JPMorgan Chase & Co. Wilmington, Delaware, United States
-
SRE Platform Engineer
GE Vernova Monterrey, Chiapas, Mexico
-
Site Reliability Software Engineer (Hybrid)
Cranial Technologies London, England, United Kingdom · $115K–$145K/yr
-
Site Reliability Engineer
Cognition San Francisco, California, United States · $260K–$300K/yr
-
Senior Site Reliability Engineer
Tempo Ireland