RocksDB WriteBatch and WAL recovery duplication
26.5K reputation · 15 Oct 2023, 17:58 UTC
RocksDB utilizes WriteBatch to ensure atomicity, ensuring that a group of operations is applied as a single unit. While this prevents partial updates, the system relies on the Write-Ahead Log (WAL) for durability and recovery after a crash.
A design uncertainty exists regarding the interaction between the WAL and the memtable during recovery. Specifically, if a WriteBatch is successfully logged to the WAL but the process terminates before the changes are fully applied to the memtable, the recovery process must replay these entries.
In scenarios where an application implements its own retry logic for failed write calls, there is a risk of duplicating writes if the previous attempt was logged but not acknowledged. Since RocksDB does not provide a native client-side transaction ID for deduplication across separate Write calls, the behavior depends on key determinism.
- How does RocksDB handle the replay of a
WriteBatchfrom the WAL to ensure no duplicate entries are created if the application retries the same batch? - Is there a mechanism within
WriteOptionsto prevent the re-application of a specific batch during recovery?
1 answer
1 question comment
Use comments to ask for clarification. Post a solution as an answer.
26,525 reputation · 15 Oct 2023, 22:41 UTC
During a write, RocksDB first assigns a monotonically increasing log number to the WriteBatch, writes the batch to the WAL (fsync if sync=true), then applies the updates to the memtable and advances the memtable’s sequence number. The MANIFEST records the highest log number that has been flushed to SST files. On restart, RocksDB replays WAL records whose log number is greater than that flushed point. If a batch was logged but the memtable update was lost because the process crashed before the apply step, its log number is still > flushed point, so recovery replays it once into a fresh memtable. After replay, the batch’s log number is now considered flushed, so a second replay would skip it because its log number ≤ flushed point. This mechanism prevents the same WriteBatch from being applied twice during recovery, regardless of client‑side retries.