Senior Site Reliability Engineer
Technology, Information and Internet · 2-10 employees
About the role
The Senior Site Reliability Engineer will design and evolve cloud infrastructure on GCP while building internal tooling to promote team autonomy. They will also advance observability platforms, manage incident response, and define reliability metrics like SLOs and DORA.
What they look for
Requirements
Candidates must have at least 3 years of SRE or production experience with proficiency in GCP, Kubernetes, and Infrastructure as Code. Strong skills in observability tools and scripting languages like Python or Go are required.
Full description
About the Role
We are a well-funded AI/ML company operating at the intersection of geospatial intelligence and environmental technology. Our engineering team is growing, and we're looking for a Senior Site Reliability Engineer to take ownership of our cloud infrastructure and help us raise the bar on reliability, observability, and operational excellence across Product & Engineering.
You'll evolve our GCP-based infrastructure, drive incident management practices, define SLOs and error budgets, and champion observability to improve mean time to recovery. You'll also use DORA metrics as a lens to help teams ship better software, and work cross-functionally to optimize cloud usage and cost.
What You'll Do
- Design and evolve cloud infrastructure on GCP at scale.
- Build internal tooling and automation that promote team autonomy and self-service.
- Advance the observability platform (metrics, logging, tracing) to reduce MTTR.
- Build visibility into infrastructure costs and drive governance and optimization initiatives.
- Champion reliability best practices including SLOs, SLIs, error budgets, and DORA metrics.
- Lead incident management, facilitate post-incident reviews, and participate in on-call rotation.
What We're Looking For
Required:
- 3+ years of Site Reliability Engineering or production SRE experience.
- Proficiency with Google Cloud Platform (GCP), including cost optimization and governance.
- Hands-on experience with Kubernetes for cluster and workload management.
- Infrastructure as Code experience using tools such as Terraform or Deployment Manager.
- Scripting and automation skills in Python, Bash, or Go.
- Strong observability stack experience: Prometheus, Grafana, OpenTelemetry, logging, and distributed tracing.
- Experience with incident management, post-incident reviews, and on-call rotation.
- Ability to define and implement SLOs, SLIs, and error budgets.
Nice to Have:
- Experience designing and evolving cloud infrastructure at scale.
- Familiarity with DORA metrics and how to apply them to engineering workflows.
- Background in AI/ML or geospatial technology environments.
Location
This role is fully remote, open to candidates based in EU, UK, or North America (Canada, United States, and select European countries including Denmark, Estonia, France, Netherlands, Portugal, Sweden, Switzerland, and the United Kingdom).
Visa sponsorship is not available.
Compensation & Benefits
Compensation details were not provided for this role. Salary will be discussed during the interview process and will be commensurate with experience and location.
Similar roles
-
Site Reliability Engineer II
Backblaze External Website Bangalore, Karnataka, India
-
Site Reliability Engineer
Firmus Technologies Singapore, Singapore
-
Site Reliability Engineer I
Backblaze External Website Bangalore, Karnataka, India
-
Staff+ Site Reliability Engineer, Safeguards ML Infra
Anthropic San Francisco, California, United States · $405K–$485K/yr
-
Linux Site Reliability Engineer (SRE)
OCBC Sepang, Selangor, Malaysia
-
Senior Lead Site Reliability Engineer
FIS pune, Maharashtra, India