Emergency Support
All News

Scaling to Petabyte Scale: New Whitepaper by SAP & CLYSO on Ceph & Rook

Learn how to operate 120 Petabytes of storage across 30 global regions in a fully digitally sovereign and automated way.

How do you operate 120 Petabytes of storage across 30 global regions in a fully digitally sovereign and automated way?

In close partnership, the SAP Engineering Team and the CLYSO Engineering Team developed a fully open-source, digitally sovereign cloud storage architecture as part of the Apeiro / ApeiroRA project.

We have summarized the insights, performance telemetry, and best practices gained from this initiative in a comprehensive whitepaper.

What is Rook and What Is It Used For?

Ceph is one of the most powerful and flexible open-source storage technologies in the world. However, operating Ceph clusters manually in a dynamic cloud environment introduces significant operational complexity.

Rook addresses this challenge: As an open-source Cloud Native Computing Foundation (CNCF) Kubernetes operator, Rook turns Ceph into an automated, declarative service. Instead of managing storage manually or executing complex scripts, the entire storage layer can be controlled directly via Kubernetes (using GitOps and Helm charts). In the background, Rook autonomously handles bootstrapping, configuration, scaling, and self-healing of the Ceph clusters.

Key Takeaways of the Whitepaper at a Glance

In our joint whitepaper, we highlight the architecture and day-to-day production operations of this global storage fleet:

  • Strict Decoupling of Hardware & Software: The interplay between the Metal-API (hardware lifecycle), Gardener (Kubernetes cluster provisioning), and Rook (Ceph orchestration).

  • Real Performance Boundaries: Results from breakpoint testing using k6 clients on Ceph Squid clusters. The whitepaper analyzes the exact deltas between client-side latencies and RGW backend latencies (connection queueing) up to the saturation limit of 90,000 GET requests/sec.

  • Predictable Operations ("Day 2 Operations"): How rolling zero-downtime upgrades (from Ceph Reef to Squid) and automated rebalancing during drive failures reliably work in live production across roughly 2,800 OSDs.

  • Digital Sovereignty: Complete independence from hyperscaler vendor lock-in, running on dedicated hardware in our own data centers.

Why You Should Read the Whitepaper

Whether you are a Platform Engineer, Infrastructure Architect, or IT Leader - if you are facing the challenge of building distributed Kubernetes storage, this document offers battle-tested engineering insights rather than abstract theory:

  • Architecture Blueprint: Utilize proven specifications and best practices for your own cloud-native projects.

  • Telemetry & Troubleshooting: Understand how latencies develop under peak load and how to realistically plan capacity.

  • Enterprise-Grade Open Source: Learn how global industry leaders and open-source maintainers collaborate on equal footing.

Read More & Get Started

Download the full whitepaper directly or dive deeper into the technical performance charts in the original community blog post:

>> Download Whitepaper (PDF)

>> Read Technical Blog Post on Medium/Rook.io

Do you have questions about cloud-native storage, Ceph, or Rook?

>> Get in touch with our engineering team!

We are happy to support you from architectural design to production operations.