Choosing a Compression Algorithm in RocksDB: A Decision Guide
A decision guide for selecting RocksDB compression algorithms (Snappy, LZ4, Zstd, etc.) with per‑column‑family tuning, a concrete C++ config example, and a reproducible db_bench micro‑benchmark to validate trade‑offs.
06 May 2026, 14:14 UTC

Problem and Takeaway
RocksDB stores data in immutable SST files that are compressed block‑by‑block. The compression algorithm you pick determines three things at once: on‑disk size, CPU spent (de)compressing, and read latency caused by decompression. Because the setting lives at the column‑family level, you can tune it per data type, but you cannot change it for existing data without a full rewrite. The practical takeaway: start with Snappy for latency‑sensitive workloads, move to Zstd when storage is the bottleneck, and validate with a realistic micro‑benchmark before committing.
Supported Compression Types (RocksDB ≥ 6.20)
| Algorithm | Typical Ratio | Compress Speed | Decompress Speed | Best Fit |
|---|---|---|---|---|
kNoCompression | 1.0× | N/A | N/A | Already‑compressed payloads (e.g., images, protobuf with snappy) |
kSnappyCompression | 1.5–2.5× | Very fast | Very fast | Low‑latency reads/writes, SSD, CPU‑constrained |
kLZ4Compression | 1.6–2.8× | Fast | Fast | Balanced speed/size, good default for mixed workloads |
kLZ4HCCompression | 2.0–3.5× | Slower (high‑compression mode) | Fast | Write‑once, read‑many data where extra CPU on write is acceptable |
kZSTDCompression | 2.5–4.5× | Moderate | Moderate | Storage‑constrained environments, HDD, archival data |
kBZip2Compression | 3.0–5.0× | Slow | Slow | Rarely used; only when maximum size reduction outweighs CPU cost |
Decision Constraints
- Hardware: SSD vs. HDD changes the I/O‑vs‑CPU trade‑off. On SSD, decompression latency is more visible.
- Workload pattern: Write‑heavy (logs, time‑series) favors faster compressors; read‑heavy (lookup tables) can tolerate slower decompression if it saves I/O.
- CPU budget: If the server runs other latency‑sensitive services, keep compression CPU low.
- Data compressibility: Already‑compressed blobs gain nothing from further compression; use
kNoCompressionfor those column families.
Concrete Configuration Example
Assume a database with two column families: metadata (small JSON blobs, high read rate) and events (large append‑only logs). The following C++ snippet shows per‑family tuning:
#include "rocksdb/options.h"
#include "rocksdb/db.h"
using namespace rocksdb;
Options db_opts;
db_opts.create_if_missing = true;
ColumnFamilyOptions meta_opts;
meta_opts.compression = kSnappyCompression; // low latency reads
meta_opts.compression_per_level.resize(7);
meta_opts.compression_per_level[0] = kNoCompression; // level‑0 often hot, keep raw
ColumnFamilyOptions events_opts;
events_opts.compression = kZSTDCompression; // maximize space savings
events_opts.compression_per_level.resize(7);
events_opts.compression_per_level[0] = kLZ4Compression; // level‑0 still fast
std::vector<ColumnFamilyDescriptor> cfd = {
{ "metadata", meta_opts },
{ "events", events_opts }
};
std::vector<ColumnFamilyHandle*> handles;
DB* db;
Status s = DB::Open(db_opts, "/mnt/rocksdb", cfd, &handles, &db);
assert(s.ok());
Key points:
compression_per_levellets you override the default for specific compaction levels (e.g., keep level‑0 uncompressed to avoid decompression on the hottest data).- Changing
compressionafter data exists only affects new SST files; old files stay with their original algorithm until they are compacted.
Validation Micro‑Benchmark
Run the built‑in db_bench tool on a staging node that mirrors production hardware. The command below writes 1 GiB of 1 KiB random keys/values with each compressor and reports write throughput, disk usage, and read latency.
# Run as the same user that owns the RocksDB data directory
for comp in snappy lz4 lz4hc zstd bzip2; do
./db_bench \
--benchmarks=fillseq,readseq \
--num=1000000 \
--value_size=1024 \
--compression_type=$comp \
--db=/tmp/rocksdb_bench_$comp \
--statistics=1 \
--threads=4
echo "--- $comp done ---"
done
What to check (run after each iteration):
- Write throughput (MB/s) – higher is better for ingest.
- Resulting DB size (bytes) – lower means better compression.
- Average read latency (µs) – watch for spikes when decompression dominates.
- CPU utilization (via
toporpidstat) – ensure it stays within your budget.
Compare the numbers against your SLAs. If Zstd saves 30 % space but adds 150 µs read latency, decide whether the storage cost outweighs the latency penalty.
Limitations and Practical Checks
- No hot‑swap: Changing compression for existing data requires a full compaction or a
CompactRangecall that rewrites all SST files. Plan a maintenance window. - Build‑time availability: Some distributions ship RocksDB without Zstd or BZip2. Verify with
ldd $(which rocksdb) | grep -E 'zstd|bz2'or by attempting to open a DB with the desired type; RocksDB will log a warning if unsupported. - Interaction with block size: Larger block sizes (default 4 KiB) improve compression ratio but increase read‑amplification for point lookups. Test block sizes 4 KiB, 16 KiB, 64 KiB together with your chosen compressor.
- Cache effects: Compressed blocks occupy less block‑cache space, potentially raising hit rates. Monitor
rocksdb.block.cache.hitandmisscounters after deployment.
Quick Verification Checklist
- Open the DB with the target compression; confirm no
Unsupported compression typewarning in LOG. - Run the
db_benchsuite on a representative dataset. - Inspect
rocksdb.statsforcompression_ratioandcompaction.read.bytesvscompaction.write.bytes. - Validate block‑cache hit rate stays ≥ 90 % for read‑heavy column families.
- Document the chosen algorithm per column family in your deployment playbook.
Summary
Pick Snappy for latency‑critical paths, Zstd when disk is scarce, and LZ4/LZ4HC as a middle ground. Apply the setting per column family, use compression_per_level to keep the hottest level fast, and always benchmark on production‑like hardware before rolling out. The configuration is static for existing data, so treat a compression change as a data‑migration event.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.