Diagnosing Pulsar Broker Ledger Segment Retention Issues
Identify and resolve Pulsar broker ledger segment retention problems that cause disk growth, cleanup warnings, or consumer lag.
11 Sept 2026, 16:54 UTC

Recognizable Condition
You notice steady disk‑space growth on a Pulsar broker, rapid increase in the number of ledger segments, or frequent SegmentCleanup warnings in the broker log. These symptoms indicate that old ledger segments are not being removed according to the configured retention policy.
Cause/Diagnostic Table
| Possible Cause | What to Look For |
|---|---|
segmentRetentionMinutes set too high | Segments persist longer than expected; disk usage grows linearly with time. |
segmentSizeMaxBytes | |
| High producer throughput on a low‑retention topic | Segment creation rate outpaces cleanup; log shows many SegmentCreated messages with low idle time. |
| BookKeeper I/O latency or cleanup thread contention | Repeated SegmentCleanup warnings/errors; cleanup thread appears stuck in thread dumps. |
| Retention policy too aggressive (very low minutes) | Consumer lag spikes while broker reports many small segments; frequent GC pressure. |
Ordered Checks
- Verify current retention settings on each broker:
pulsarctl broker info --segment-retention --broker : - Scan the broker log for cleanup activity:
grep SegmentCleanup /var/log/pulsar/broker.log - Check segment size and count for a topic:
pulsarctl topics stats --tenant --namespace --topic - Monitor disk usage trend:
df -h /path/to/bookkeeper/journal(run at intervals). - If BookKeeper latency is suspected, review BookKeeper server logs for
AddEntrylatency spikes.
Fixes Tied to Findings
- Retention minutes too high – Decrease
segmentRetentionMinutesinbroker.conf(e.g., from 1440 to 720). After editing, restart the broker to activate the new cleanup interval. - Segment size too large – Lower
segmentSizeMaxBytes(e.g., from 100MB to 10MB). This forces more frequent segment rolls, giving the cleanup thread more opportunities to delete old data. - Throughput outpaces cleanup – Increase
segmentRetentionMinutesslightly or add more broker instances to share load; alternatively, enable topic‑level retention overrides if supported. - Cleanup thread stalled – Investigate BookKeeper I/O: check disk performance, network latency, and increase
bookkeeperClientTimeoutMillisif needed. Ensure the cleanup thread pool size (segmentCleanupThreadPoolSize) is adequate. - Over‑aggressive retention causing lag – Raise
segmentRetentionMinutesto a value that matches your replay window (e.g., 30 minutes) and monitor consumer lag.
Escalation Criteria
- Disk usage continues to grow > 80 % of available space after applying the above fixes.
- Broker logs show persistent
SegmentCleanuperrors despite healthy BookKeeper performance. - Consumer lag remains high (> 5 minutes) and segment count does not decrease after configuration adjustment.
- Multiple brokers exhibit the same symptom, indicating a cluster‑wide configuration issue.
Limitation: Changing retention values may affect data replay capabilities; always verify that the new window satisfies your application’s replay or audit requirements before rolling out to production.
Practical verification: After a broker restart, re‑run pulsarctl broker info --segment-retention to confirm the new values, then watch the broker log for SegmentCleanup messages indicating successful deletion. Segment count for a topic should trend downward over the next few retention intervals.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.