gridscale GmbH

Site Reliability Engineer (m/f/d)

gridscale GmbH · Cologne, North Rhine-Westphalia, Germany

Technology, Information and Internet · 51-200 employees

19 h ago
Mid (2-5 yrs) Full-time Germany
Log in to apply, save this posting, or score it against your profile with AI.

About the role

You will build and maintain the framework for deploying Cloud Store packages while driving the development of the Kubernetes stack and GitOps workflows. Additionally, you will support other teams with onboarding and participate in system analysis to implement continuous improvements.

What they look for

Kubernetes GitOps Ansible Terraform Prometheus Grafana Linux OpenStack Python Bash Go CI/CD Infrastructure as Code Observability Security fundamentals Distributed systems

Requirements

The role requires deep knowledge of Linux system administration and hands-on experience with cloud platforms like OpenStack and Kubernetes. You must also be proficient in Infrastructure as Code tools, CI/CD pipelines, and scripting languages such as Python, Bash, and Go.

Benefits

32 vacation days Flexible working hours Home office options Employer-funded pension plan Insurance package Public transportation subsidy Sports activities contribution Corporate discounts Cargo bike leasing Company events Free beverages

Full description

At our company, it’s all about #OneTeam! Join gridscale and help shape the future of the cloud together with OVH.

As a leading tech company, we’ve been working for over two decades to reduce our environmental footprint - with innovative solutions and an open cloud designed to be sustainable from the ground up: #SustainableByDesign.

As the Cloud Store team, we provide the framework to automate an standardized deployment on top of the OnPrem Cloud Platform. Are you passionate about air-gapped cloud environments and edge technologies? Then you've come to the right place.

Our Tech Stack 🚀

•Kubernetes •GitOps •Ansible •Terraform •Prometheus

•Grafana •Linux •OpenStack•Python, Bash & Go

Your Role💻

As a Site Reliability Engineer, you'll be part of a team responsible for building the framework to package and deploy software solutions on top of OPCP. You'll work on the conception, automation, and operation of our platform and drive continuous improvements. You'll help new teams onboard with the framework and OPCP. We're looking for someone who is comfortable in a security-oriented environment with a high degree of automation (GitOps). You are used to maintaining an overview of ambiguous situations, analyzing systems, and making well-founded decisions based on this analysis

Your Tasks

  • Build the framework for building and deploying Cloud Store packages
  • Drive the ongoing development of our Kubernetes stack and implement GitOps workflows (e.g., with FluxCD)
  • Support other teams and help them onboard with OPCP
  • Actively participate in system analysis and derive improvements, even when initial requirements are unclear
  • Develop and maintain our Infrastructure as Code using tools like Ansible and Terraform
  • Participate in an on-call rotation

What we offer you💼

  • Exceptional team spirit across all departments and national borders; we live #OneTeam
  • Exciting work in a highly innovative and international environment with cutting-edge technologies
  • 32 vacation days, increasing with length of service
  • Flexible working hours, home office options, and a secure permanent position with market- and performance-based compensation
  • Employer-funded pension plan and an attractive insurance package
  • OVHcloud covers 50% of public transportation costs
  • Up to €400 annual financial contribution from OVHcloud towards sports activities (gym membership, sports classes, etc.)
  • Through Corporate Benefits, you receive attractive discounts at numerous shops and companies
  • We contribute to the leasing of your cargo bike
  • Regular company events and free cold and hot beverages
  • Deep knowledge of Linux/Unix system administration and internals
  • Hands-on experience with cloud platforms, especially OpenStack, including infrastructure provisioning and management
  • Proficiency with Infrastructure as Code tools such as Terraform and Ansible
  • Experience with containerization and orchestration technologies, especially deep Kubernetes
  • Skilled in building and maintaining CI/CD pipelines
  • Hands-on experience with modern observability stacks, including metrics, logs and traces
  • Proficiency in scripting and automation using Python, Bash, and Go
  • Understanding of security fundamentals including IAM, secrets management, hardening and compliance
  • Understanding of distributed systems concepts such as the CAP theorem, consensus and fault tolerance
  • Strong problem-solving skills, adaptability and strict documentation discipline
  • Ability to thrive in collaborative product-team environments with a strong ownership mentality and a blameless culture mindset
  • Strong communication skills and a passion for continuous improvement