Write Stall Mode: kStopWritesFirst vs kSlowdown for p99 Latency Under Concurrent Writers
0 reputation · 30 May 2021, 20:16 UTC
0 reputation · 30 May 2021, 20:16 UTC
Bound p99 write latency when multiple threads ingest into RocksDB concurrently, without sacrificing sustained throughput.
RocksDB triggers a write stall when memtable usage, L0 file count, or compaction debt exceed internal thresholds. Under concurrent writers this stall manifests as sudden p99 spikes. Two documented stall modes offer different backpressure semantics:
Status::TryAgain once the stall limit is reached, providing hard latency bounds at the cost of client-side retries.The optimal choice depends on workload burstiness, retry logic availability, and whether the service can tolerate rejected requests versus elongated pauses. Thresholds (write_stall_limit, max_write_buffer_number, level0_stop_writes_trigger) interact with the selected mode in ways that are not fully deterministic across hardware and I/O schedulers.
write_stall_limit be tuned relative to the rate limiter’s bytes_per_sync and delayed_write_rate when switching between modes?A thoughtful contribution can make all the difference. Be the first to share one.
Use comments to ask for clarification. Post a solution as an answer.
No question comments on this page.