Architecting Long-Term Retention with Apache Pulsar Tiered Storage
Learn how to implement Apache Pulsar Tiered Storage to decouple data retention from compute costs while maintaining transparent data access via object stores.
10 Oct 2025, 01:13 UTC

The Storage Cost Dilemma in High-Throughput Messaging
Scaling data retention in a standard distributed log usually requires adding more disks to the storage layer (BookKeeper in Pulsar’s case). This creates a linear cost increase: to keep data for a year instead of a week, you must scale your most expensive, high‑performance disks, even if 99% of your reads target data from the last hour. The goal is to decouple the cost of storage from the performance of the real‑time write path.
The Smallest Suitable Design
The most efficient implementation of long‑term retention is a Tiered Storage architecture. Instead of scaling the BookKeeper cluster, you introduce an object store (S3, GCS, or Azure Blob Storage) as a secondary storage tier.
In this design, the Pulsar Broker acts as the orchestrator. It monitors ledgers (the segments of data stored in BookKeeper) and offloads them to the object store once they meet specific criteria, such as age or size. Crucially, the Broker maintains the metadata, so consumers remain unaware of whether a message is being served from a local disk or a remote cloud bucket.
Trust and Data Boundaries
When implementing tiered storage, the data boundary shifts from a closed cluster to an external service. This introduces three critical considerations:
- Access Control: The Broker requires IAM roles or secret keys to write to the object store. These credentials should be scoped to the specific bucket used for offloading to prevent lateral movement.
- Network Transit: Data moving from BookKeeper to S3 travels over the network. In multi‑region setups, ensure the object store bucket is in the same region as the Pulsar cluster to avoid egress costs and latency spikes.
- Consistency: The system relies on the object store’s eventual or strong consistency. Since offloaded ledgers are immutable (they are only written once and never changed), the risk of consistency conflicts is low.
Operational Configuration
To enable tiered storage, you must configure the broker to use a specific offload driver. Below is a representative configuration for an AWS S3 environment in broker.conf.
# Enable tiered storage
managedLedgerOffloadDriver=amazon-s3
# S3 Specific Configuration
# Note: Ensure the Broker process has the necessary IAM permissions
# or provide explicit credentials via environment variables
managedLedgerOffloadS3Bucket=pulsar-retention-bucket
managedLedgerOffloadS3Region=us-east-1
Verification Steps:
- Run the Pulsar admin CLI from a terminal with
pulsar-adminaccess to trigger a manual offload for a specific topic:pulsar-admin topics offload-ledger // - Check the object store console to confirm the creation of the ledger files.
- Consume a message with a cursor set to a timestamp preceding the offload threshold to verify transparent retrieval.
Failure Modes and Latency Trade‑offs
| Failure/Constraint | Impact | Mitigation |
|---|---|---|
| Object Store Downtime | Unable to read "cold" data; real‑time writes unaffected. | Use regionally replicated buckets (e.g., S3 Cross‑Region Replication). |
| Read Latency Spike | Significant increase in TTFB (Time to First Byte) for old data. | Implement a read‑ahead cache at the broker level. |
| Premature Offloading | Frequent S3 requests for "warm" data, increasing costs. | Tune the offloadThreshold based on actual consumer lag patterns. |
Conditions for Redesigning the Approach
Tiered storage is the correct choice when the read pattern is heavily skewed toward recent data. However, you should reconsider this architecture if:
- Random Access Patterns: If your application frequently performs random reads across the entire history of the topic, the latency of object storage will become a bottleneck.
- Strict Latency SLAs: If "cold" data must be retrieved with sub‑millisecond latency, you must scale BookKeeper horizontally with NVMe drives rather than offloading.
- Data Sovereignty: If regulatory requirements forbid storing data in a shared object store environment, you may need to implement a custom offload driver targeting a private, on‑premises S3‑compatible store (e.g., MinIO).
Rollback: Because offloading is a metadata operation that copies data rather than moving it (until the BookKeeper ledger is deleted by retention policies), the "rollback" is simply disabling the offload driver in broker.conf and restarting the brokers. This stops new offloads but does not delete existing cloud data.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.