Berge Group

TSP SRE expert

Berge Group Gothenburg, Nebraska, United States

Design Services · 51-200 employees

5 h ago
sre Senior (5-10 yrs) Full-time United States
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

The role involves managing end-to-end observability, multi-cloud FinOps, and production reliability governance. You will design and operate unified dashboards, alerting frameworks, and optimize Kubernetes platforms across AWS and GCP.

What they look for

SRE Kubernetes AWS GCP Terraform Pulumi CDK Prometheus Grafana Loki Elastic Kibana Java Python Go FinOps

Requirements

Candidates must have over 5 years of experience in DevOps or SRE, with at least 3 years in a Staff or Principal engineering capacity. Proficiency in cloud-native technologies, infrastructure as code, and programming languages like Java, Python, or Go is required.

Full description

🧩 ABOUT ENGAGE STUDIOS Engage Studios is a Swedish partner for creative, engineering and production services. We connect skilled specialists with exciting assignments for clients developing products and experiences of tomorrow. 💼 WE OFFER - Longterm, full-time TSP SRE expert assignment - Ownership of observability, multi-cloud FinOps and production reliability - Staff/Principal-level work in cloud and Kubernetes 🧪 JOB DESCRIPTION We seek a TSP SRE expert for end-to-end observability, multi-cloud FinOps, production reliability governance and Kubernetes platform optimisation. Design and operate unified observability, dashboards, alerting, SLI/SLO/SLA and error-budget frameworks, AWS/GCP cost optimisation, incident governance, and EKS/GKE optimisation. On-call is approximately one week per month. ✅ REQUIREMENTS - 5+ years DevOps, SRE or cloud-platform experience - 3+ years Staff/Principal engineer or Tech Lead ownership - Large-scale distributed systems in production - Deep hands-on AWS and GCP, including cross-cloud architecture - Terraform, Pulumi or CDK infrastructure as code - Prometheus, Grafana, Loki, Elastic and Kibana - Deep Kubernetes internals and production clusters - Java or Python/Go programming - Reliability governance, RCA, gradual rollout, rollback and SLO practices - Full-time longterm assignment, on-call, and onsite west coast of Sweden 👤 YOUR PROFILE Senior SRE or platform engineer who leads by doing and owns production outcomes. Google SRE, chaos engineering, eBPF, multi-cloud DR or Mandarin Chinese are advantages. 📍 PLACE OF EMPLOYMENT Onsite on the west coast of Sweden. No travel required. ℹ️ OTHER Full-time, 100%. On-call approximately one week per month. Applications must be submitted through this link and code 1322

Similar roles