Senior Site Reliability Engineer
Jobgether United States
Internet Marketplace Platforms · 11-50 employees
About the role
You will design and evolve scalable cloud infrastructure on Google Cloud Platform while enabling developers through automation and self-service tooling. Additionally, you will champion modern reliability practices, lead incident management, and optimize cloud efficiency.
What they look for
Requirements
Candidates must have 3+ years of professional experience in Site Reliability Engineering or a related infrastructure role. Strong hands-on experience with Google Cloud Platform, Kubernetes, and Infrastructure as Code tools like Terraform is required.
Benefits
Full description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer based in United States.
Join a growing AI/ML organization operating at the intersection of geospatial intelligence and environmental technology. As a Senior Site Reliability Engineer, you will take ownership of cloud infrastructure and help strengthen reliability, observability, and operational excellence across engineering and product teams. You will design and evolve scalable infrastructure on Google Cloud Platform while enabling developers through automation and self-service tooling. Your work will directly influence system availability, incident response, deployment performance, and cloud efficiency. You will champion modern reliability practices, from SLOs and error budgets to DORA metrics and observability. This is a fully remote opportunity offering significant technical ownership in a collaborative, high-impact engineering environment.
\n
Accountabilities
- Design, build, and continuously evolve scalable and reliable cloud infrastructure on Google Cloud Platform.
- Develop internal tooling, automation, and self-service capabilities that increase engineering efficiency and reduce operational dependencies.
- Strengthen the observability platform across metrics, logging, distributed tracing, and monitoring to improve system visibility and reduce mean time to recovery.
- Establish infrastructure cost visibility, governance, and optimization initiatives to improve cloud efficiency and manage spending responsibly.
- Define and champion reliability practices including Service Level Objectives (SLOs), Service Level Indicators (SLIs), error budgets, and DORA metrics.
- Lead incident management activities, coordinate response efforts, facilitate post-incident reviews, and drive meaningful improvements based on incident learnings.
- Participate in the on-call rotation and help ensure production systems remain stable, available, and resilient.
- Partner with Product and Engineering teams to improve operational practices, deployment reliability, and overall software delivery performance.
Requirements
- 3+ years of professional experience in Site Reliability Engineering, production engineering, or a closely related infrastructure role.
- Strong hands-on experience with Google Cloud Platform, including cloud cost optimization, governance, and infrastructure management.
- Proven experience managing Kubernetes clusters and workloads in production environments.
- Experience with Infrastructure as Code tools such as Terraform or Google Cloud Deployment Manager.
- Strong scripting and automation capabilities using Python, Bash, Go, or comparable languages.
- Extensive experience with observability technologies such as Prometheus, Grafana, OpenTelemetry, centralized logging, and distributed tracing.
- Practical experience with incident management, post-incident analysis, production troubleshooting, and on-call operations.
- Ability to define, implement, and monitor SLOs, SLIs, and error budgets.
- Strong understanding of reliability engineering principles and a proactive approach to identifying and resolving operational risks.
- Familiarity with DORA metrics and their application to engineering workflows is an asset.
- Experience working in AI/ML, geospatial technology, or other data-intensive technical environments is considered a plus.
- Strong communication and collaboration skills, with the ability to work effectively across distributed Product and Engineering teams.
- Candidates must be authorized to work in their country of residence; visa sponsorship is not available.
Benefits
- Fully remote work environment.
- Opportunity to work with modern cloud infrastructure, Kubernetes, Infrastructure as Code, and advanced observability technologies.
- Significant ownership over infrastructure reliability, operational excellence, and cloud optimization initiatives.
- Opportunity to influence engineering practices through SLOs, error budgets, DORA metrics, and incident management.
- Collaboration with cross-functional and distributed engineering teams across North America, Europe, and the UK.
- Exposure to innovative AI/ML and geospatial technology applications.
- Compensation will be discussed during the interview process and will be aligned with experience and geographic location.
\nHow Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1
Similar roles
-
Site Reliability Engineer II
Backblaze External Website Bangalore, Karnataka, India
-
Site Reliability Engineer
Firmus Technologies Singapore, Singapore
-
Site Reliability Engineer I
Backblaze External Website Bangalore, Karnataka, India
-
Staff+ Site Reliability Engineer, Safeguards ML Infra
Anthropic San Francisco, California, United States · $405K–$485K/yr
-
Linux Site Reliability Engineer (SRE)
OCBC Sepang, Selangor, Malaysia
-
Senior Lead Site Reliability Engineer
FIS pune, Maharashtra, India