How does RocksDB ensure that retrying a WriteBatch after a write stall does not create duplicate writes in the WAL?
0 reputation · 19 Oct 2023, 08:56 UTC
0 reputation · 19 Oct 2023, 08:56 UTC
Determine whether RocksDB's internal retry loop for WriteBatch operations guarantees that no duplicate writes are persisted to the WAL or applied to the memtable when a write stall (e.g., memtable full) triggers a retry.
Assume default WAL enabled, sequence numbers are incremented on each retry, and recovery replays only the highest sequence number per key. Consider the impact of disabling WAL, using non‑idempotent merge operators, or custom write‑callback logic on this guarantee.
What conditions, if any, allow duplicate WAL entries to survive recovery? Does disabling WAL remove the safety net entirely? How do non‑idempotent merge operators behave during repeated retries of the same WriteBatch?
29775 reputation · 19 Oct 2023, 17:41 UTC
A write stall in RocksDB is not a partial failure — it is a block before the write happens. When memtables fill up or compaction falls behind, the write thread is made to wait (or is stopped) before the WriteBatch is assigned a sequence number, appended to the WAL, or applied to the memtable. Because nothing was applied, retrying after the stall clears cannot create a duplicate. The stall path and the duplicate path are different paths; duplicates require an ambiguous failure, not a stall.
Each successful Write() goes through a single critical section in which the batch is assigned one sequence number and appended to the WAL as one logical record, then inserted into the memtable. The batch is the atomic unit: either the whole record lands in the WAL or none of it does. During recovery, WAL replay applies records in sequence-number order and skips anything at or below the last flushed sequence number, so even a WAL entry whose effects already reached an SST is not applied twice. You can confirm this on your build by checking db->GetLatestSequenceNumber() before and after a write — a successful batch consumes exactly one sequence number (plus one per key in the batch, but a single contiguous allocation).
The likely point of confusion: RocksDB does not have an internal retry loop that re-issues your batch after a stall. The stall is a wait inside the original call. The call returns Status::OK() once, meaning applied once. The real duplication risk is entirely on the caller side:
Write() of the same batch is a new batch with a new sequence number and a second WAL record. Both survive recovery. Key retries off the returned Status, always.Duplicate WAL entries surviving recovery: only if the batch was genuinely written twice (blind application retry) or if an ambiguous failure was retried and both attempts landed. Recovery itself never duplicates a single record.
Disabling WAL: it removes the recovery mechanism, not a dedup mechanism — there never was one. With WAL off, a crash loses everything since the last flush, and a retry after a crash-restart is safe only because the original write is gone. Stall behavior is unchanged: still a wait, still no partial application.
Non-idempotent merge operators: this is the sharpest edge. A merge like "increment counter" applied twice yields +2. Since a stall never causes double application, merges are safe across stalls — but an application-level blind retry of a merge batch double-applies it, and recovery will faithfully replay both records. Merge operands are not deduplicated by key; each record replays.
Stall trigger names and status subcodes vary across RocksDB versions, so confirm against yours. A practical check: enable INFO-level logging, fill the memtables until you see Stalling writes because of ... messages, issue the WriteBatch, verify it returns OK exactly once, then reopen the DB and count applied keys. If the count matches exactly-once application, the stall path behaved as described. If you need crash-safe retries with ambiguous failures, evaluate TransactionDB plus an application-level idempotency key rather than relying on engine behavior.
Use comments to ask for clarification. Post a solution as an answer.
No question comments on this page.