Globaldev Group

Senior Ceph/Rook-Ceph + Kubernetes Storage Engineer

Globaldev Group Portugal

IT Services and IT Consulting · 201-500 employees

6 h ago
Remote kubernetes Senior (5-10 yrs) Full-time Portugal
Log in to apply, save this posting, or score it against your profile with AI.

About the role

The role involves managing and troubleshooting production Ceph environments and Kubernetes storage infrastructure. You will be responsible for scaling, performance tuning, and ensuring high availability for business-critical storage systems.

What they look for

Ceph Rook-Ceph Kubernetes DevOps SRE Infrastructure Engineering Linux Administration Ansible Helm Prometheus Grafana CSI Storage Architecture Performance Tuning Troubleshooting Automation

Requirements

Candidates must have over 5 years of experience in infrastructure or storage engineering with deep hands-on knowledge of Ceph and Kubernetes. Strong skills in Linux administration, automation with Ansible, and observability tools like Prometheus and Grafana are required.

Full description

We’re currently looking for a Senior Ceph/Rook-Ceph + Kubernetes Storage Engineer to join a long-term project. The role focuses on production Ceph environments, Kubernetes storage, and infrastructure troubleshooting.

  • 5+ years in DevOps, SRE, infrastructure or storage engineering
  • Deep hands-on production Ceph experience: OSD, MON, MGR, PGs, recovery, backfill, capacity planning, performance tuning, scaling and high availability
  • Experience building, operating, upgrading and troubleshooting production Ceph clusters
  • Hands-on Rook-Ceph in Kubernetes, preferably business-critical production environments
  • Strong Kubernetes knowledge, including CSI/storage integration
  • Linux administration and troubleshooting at OS/hardware level
  • Automation experience with Ansible
  • Kubernetes tooling such as Helm
  • Monitoring/observability with Prometheus and Grafana
  • Experience investigating storage performance, latency, disk failures, network bottlenecks and recovery issues
  • Strong understanding of failure domains, CRUSH topology, replication and storage architecture
  • Good English communication skills

Highly valuable but not mandatory:

  • OpenStack experience, especially Ceph integration with Cinder, Glance, Nova and RBD
  • Petabyte-scale Ceph environments
  • Bare-metal infrastructure
  • Large production Rook-Ceph clusters
  • CephFS and RGW/S3 experience
  • Customer-facing troubleshooting/support experience

Similar roles