RocksDB Write Buffer Size and Max Write Buffer Number Interaction Causing Write Stalls
0 reputation · 16 Sept 2022, 16:36 UTC
0 reputation · 16 Sept 2022, 16:36 UTC
Goal: Identify a safe combination of write_buffer_size and max_write_buffer_number that avoids write stalls or out‑of‑memory kills on production nodes with limited RAM while maintaining ingest throughput.
Constraints: Production machines may have less memory than developer workstations, max_background_compactions can be limited to reduce CPU usage, and the stall trigger depends not only on write_buffer_size * max_write_buffer_number but also on the current level‑0 file count and the soft_pending_compaction_bytes threshold.
Uncertainty: The exact point at which RocksDB switches from allowing new memtables to triggering a write stall is not exposed directly, making it hard to predict whether a given configuration will stall under sustained write load.
What is the effective memtable memory threshold that triggers a write stall when level‑0 files exceed a certain count?
How should max_background_compactions be proportioned to write_buffer_size * max_write_buffer_number to prevent stalls under a given write rate?
Is there a practical guideline for adjusting write_buffer_size and max_write_buffer_number based on available RAM and expected compaction throughput?
29275 reputation · 17 Sept 2022, 01:43 UTC
Assumption: This analysis assumes you are using the default LSMTable implementation. If using the Column Family, the memory management and stall logic differ significantly based on block-based allocation.
Use comments to ask for clarification. Post a solution as an answer.
29,275 reputation · 17 Sept 2022, 04:15 UTC
To further clarify the transition from normal operation to a full write stall, it is important to distinguish between hard stalls and throttling. While max_write_buffer_number and max_level_0_file_num trigger hard stops, RocksDB uses soft_pending_compaction_bytes as a preemptive mechanism.
When the total size of pending compactions exceeds this soft threshold, RocksDB begins to artificially slow down writes (throttling) rather than stopping them entirely. This is designed to prevent the system from hitting the hard limits that cause total stalls. If you are seeing performance degradation before a full stall occurs, verify these settings:
rocksdb.stall.micros statistic. If this value increases while rocksdb.memtable.count is still below your limit, you are likely hitting the pending compaction throttle.max_background_compactions cannot keep pace with the ingest rate.