Diagnosing and Resolving RocksDB Write Stall Conditions
Learn how to spot RocksDB write stalls, diagnose their root causes with specific properties, apply targeted fixes, and know when to escalate.
29 Aug 2025, 19:57 UTC

Recognizable condition
A write stall is observable when client‑side write latency consistently exceeds a typical threshold (e.g., >10 ms) and the RocksDB log contains messages such as "Write stall" or "stall condition".
Cause / diagnostic indicator table
| Cause | What to look for |
|---|---|
| Too many level‑0 files (exceeds max‑write‑buffer‑number or level0‑stop‑writes‑trigger) | High rocksdb.num-files-at-level0 value |
| Memtable flush lag (insufficient background flush threads) | Rising rocksdb.memtable-flush-pending |
| Compaction falling behind (max‑background‑compactions too low) | Increasing rocksdb.num-running-compactions or stalled compaction queue |
| Block cache starvation (read‑induced write stalls) | Low cache hit ratio, high rocksdb.block-cache-miss |
| Disk I/O saturation or high latency storage | Elevated rocksdb.io-rate or high %util from iostat |
| WAL sync delays (fsync overhead) | Frequent rocksdb.wal-sync-delay spikes |
Ordered diagnostic checks
- Enable RocksDB statistics (if not already on). In your DB opening code set
DBOptions::statistics = rocksdb::CreateDBStatistics();. After the DB is open, periodically calldb->GetProperty("rocksdb.write-stall..."); a non‑zero value indicates a stall. - Check level‑0 file count: run
db->GetProperty("rocksdb.num-files-at-level0"). Compare the result to the configuredlevel0-stop-writes-trigger(default varies by version). - Review memtable flush queue: query
db->GetProperty("rocksdb.memtable-flush-pending"). A steadily growing number signals flush lag. - Examine compaction queue depth: read
db->GetProperty("rocksdb.num-running-compactions")anddb->GetProperty("rocksdb.compaction-pending"). High pending compactions suggest the compaction thread pool is insufficient. - Assess storage latency: on the host run
iostat -x 1(requires read access to block devices; typically root or a user in thediskgroup). Look for high%utilor await >10 ms. Inside RocksDB you can also checkdb->GetProperty("rocksdb.io-rate").
Fixes tied to findings
High level‑0 file count
Increase the write‑buffer‑number or raise the stall triggers:
Options opts;
opts.set_max_write_buffer_number(4); // default 2
opts.set_level0_slowdown_writes_trigger(20); // default 20
opts.set_level0_stop_writes_trigger(30); // default 30
After applying, restart the instance and verify that rocksdb.num-files-at-level0 stays below the stop‑writes trigger.
Memtable flush lag
Add more background flush threads or enlarge the memtable size:
opts.set_max_background_flushes(4); // default 1
opts.set_write_buffer_size(256 << 20); // 256 MiB, default 64 MiB
Monitor rocksdb.memtable-flush-pending; it should trend toward zero.
Compaction lag
Scale up compaction threads and ensure I/O bandwidth:
opts.set_max_background_compactions(6); // default 2
// If using a dedicated SSD, verify that the device can sustain the required MB/s.
Check that rocksdb.num-running-compactions does not consistently hit the max limit and that rocksdb.compaction-pending declines.
Block cache starvation
Enlarge the block cache:
opts.set_block_cache(rocksdb::NewLRUCache(2 << 30)); // 2 GiB cache
Verify cache hit ratio via rocksdb.block-cache-hit vs rocksdb.block-cache-miss.
WAL sync delays
Enable group commits or move the WAL to faster media:
opts.set_wal_ttl_seconds(0); // disable TTL if not needed
opts.set_wal_size_limit_MB(0); // let RocksDB manage size
// Alternatively, set opts.db_log_dir to a separate NVMe device.
Watch rocksdb.wal-sync-delay for reduction.
Escalation criteria
- If after applying the above adjustments the write‑latency remains >10 ms for more than 5 minutes and
rocksdb.num-files-at-level0continues to grow beyond the stop‑writes trigger, consider a deeper storage review (firmware, queue depth, or NVMe namespace health). - When CPU utilization is near 100 % while increasing background threads, further thread adds may cause contention; at this point evaluate workload sharding or upgrading to a higher‑core‑count host.
- Persistent WAL sync delays despite moving the log to faster storage may indicate a kernel‑level fsync bottleneck; engage the OS/storage team to examine
syncsystem call latency.
All changes should be made in a staging environment first, validated with the verification steps below, then rolled out to production.
Verification
- After a configuration change, restart the RocksDB instance.
- Confirm stall absence:
db->GetProperty("rocksdb.write-stall...")returns 0. - Ensure the relevant metric stays within safe bounds (e.g., level‑0 files < stop‑writes trigger, memtable‑flush‑pending ≈ 0).
- Measure write latency before and after using
db->GetProperty("rocksdb.db-write-latency")or an external histogram; a sustained drop below the threshold indicates the stall is resolved.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.