Short answer: RocksDB does not deduplicate retried batches for you, but it also doesn't need to — as long as your writes are idempotent. WAL replay re-applies a logged batch exactly once, and a client retry of the same batch simply writes the same keys again. For plain Put/Delete, re-applying the same value produces the same state, so no duplicate entries appear. Duplication only becomes real for non-idempotent operations like accumulating Merge operators or counters you maintain outside RocksDB.
What actually happens during recovery
Each WriteBatch is appended to the WAL as a single record and tagged with a sequence number. On reopen, RocksDB replays WAL entries whose sequence numbers are above the last persisted point (the state captured in the MANIFEST and flushed SST files). A batch that was logged but never made it into a flushed memtable is replayed once, into a fresh memtable. There is no path where replay itself applies one WAL record twice.
The scenario you describe — batch logged, process killed before the write call returned — means the client never saw an acknowledgement. If the client retries, RocksDB treats that as a brand-new write with a new sequence number. For deterministic keys and values, the second write overwrites the first and the final state is identical. The "duplicate" exists only as two historical versions internally, which compaction eventually collapses; it is never visible through reads.
Where duplication is real
- Merge operators that accumulate (e.g., append-to-list or add-to-counter merges): applying the same merge twice changes the result. Do not assume your merge operator is replay-safe — verify it under duplicate application before relying on retries.
- Side state outside RocksDB: if you track offsets, counters, or queue positions in a separate store and replay them after a crash, RocksDB's atomicity can't help you.
- Lost writes misread as duplicates: with
WriteOptions.sync = false or disableWAL = true, a crash can lose acknowledged writes, and re-running your producer may create confusing state. That's a durability gap, not a replay bug.
Is there a WriteOptions flag to prevent re-application?
No. WriteOptions controls durability and latency (sync, disableWAL, no_slowdown), not deduplication. RocksDB has no client-supplied transaction ID that would let recovery skip a previously applied batch. The supported patterns are:
- Make writes idempotent: use deterministic keys so retries overwrite rather than append.
- Commit offsets atomically: in queue- or Kafka-style ingestion, store the source offset in the same
WriteBatch as the data. Both commit or neither does, so a retry can detect "already applied" by reading the offset first.
- Use TransactionDB if you need commit/abort semantics across keys, though the underlying WAL replay rule is unchanged.
Verify it yourself
Behavior here is option- and version-sensitive (e.g., manual_wal_flush, wal_bytes_per_sync defaults vary by release), so confirm against your version with a crash-injection test:
// 1. Write batch with sync=true, kill -9 the process before flush
// 2. Reopen the DB and compare key/values against expected state
// 3. Retry the same batch, reopen again — state should be unchanged
// for Put/Delete workloads
If after this test you observe actual value-level differences with plain Puts, the one diagnostic detail worth checking is whether a custom MergeOperator or a secondary index maintained by your application is involved — that's where the duplication almost always lives.