Globaldev Group

Senior Ceph/Rook-Ceph + Kubernetes Storage Engineer

Globaldev Group Romania

IT Services and IT Consulting · 201-500 employees

Aug 11
Remote kubernetes Senior (5-10 yrs) Full-time Romania
Log in to apply, save this posting, or score it against your profile with AI.

About the role

The role involves managing production Ceph environments and Kubernetes storage infrastructure. You will be responsible for troubleshooting, performance tuning, scaling, and ensuring high availability of storage clusters.

What they look for

Ceph Rook-Ceph Kubernetes Infrastructure Troubleshooting DevOps SRE Storage Engineering Ansible Helm Prometheus Grafana Linux Administration CSI Performance Tuning High Availability Capacity Planning

Requirements

Candidates must have over 5 years of experience in DevOps, SRE, or storage engineering with deep hands-on production Ceph and Kubernetes expertise. Strong skills in Linux administration, automation with Ansible, and monitoring tools like Prometheus and Grafana are required.

Full description

We’re currently looking for a Senior Ceph/Rook-Ceph + Kubernetes Storage Engineer to join a long-term project. The role focuses on production Ceph environments, Kubernetes storage, and infrastructure troubleshooting.

  • 5+ years in DevOps, SRE, infrastructure or storage engineering
  • Deep hands-on production Ceph experience: OSD, MON, MGR, PGs, recovery, backfill, capacity planning, performance tuning, scaling and high availability
  • Experience building, operating, upgrading and troubleshooting production Ceph clusters
  • Hands-on Rook-Ceph in Kubernetes, preferably business-critical production environments
  • Strong Kubernetes knowledge, including CSI/storage integration
  • Linux administration and troubleshooting at OS/hardware level
  • Automation experience with Ansible
  • Kubernetes tooling such as Helm
  • Monitoring/observability with Prometheus and Grafana
  • Experience investigating storage performance, latency, disk failures, network bottlenecks and recovery issues
  • Strong understanding of failure domains, CRUSH topology, replication and storage architecture
  • Good English communication skills

Highly valuable but not mandatory:

  • OpenStack experience, especially Ceph integration with Cinder, Glance, Nova and RBD
  • Petabyte-scale Ceph environments
  • Bare-metal infrastructure
  • Large production Rook-Ceph clusters
  • CephFS and RGW/S3 experience
  • Customer-facing troubleshooting/support experience

Similar roles