Kafka Log Compaction: Keep Only the Latest State Per Key and When to Use It
When you need to replay only the most recent value for each key in a Kafka topic, log compaction is the right tool. This guide explains how compaction works, shows a practical example, and discusses trade‑offs and limits so you can decide if it fits your use case.
05 Sept 2025, 22:19 UTC

Why Kafka Log Compaction Matters
In many event‑driven systems you only care about the current state of an entity, not every change that happened over time. Storing every update can bloat disk usage and slow down consumers that need to reconstruct the latest state. Kafka’s log compaction feature solves this by keeping only the most recent record for each key while still allowing the topic to act as a log for new events.
What Log Compaction Does
When a topic’s cleanup.policy is set to compact, the broker asynchronously scans log segments and removes older records that share the same key. The compaction algorithm keeps the newest record for each key in the retained segments, discarding older duplicates. Records without a key are never compacted.
Compaction is separate from the normal retention.ms or retention.bytes policies. A record may survive the retention window until a compaction cycle runs, so you can still see old data if compaction is delayed.
When to Use Compaction
- State stores – Rebuilding a key/value store from the topic (e.g., Kafka Streams state stores).
- Configuration topics – Storing the latest configuration for services where only the newest change matters.
- Idempotent event streams – Ensuring that consuming the topic yields the latest value without replaying all updates.
Key Configuration Parameters
| Parameter | Typical Value | Effect |
|---|---|---|
| cleanup.policy | compact | Enables compaction. |
| segment.bytes | 1073741824 (1 GB) | Size of a log segment; smaller segments compact faster but increase overhead. |
| min.cleanable.dirty.ratio | 0.5 | Minimum ratio of dirty (uncompacted) data to segment size before compaction starts. |
| max.compaction.lag.ms | 86400000 (24 h) | Maximum age of a record before it becomes eligible for compaction. |
| delete.retention.ms | 86400000 | Grace period after a tombstone before the key is removed. |
Concrete Example: Enabling Compaction on a Topic
Create a compacted topic
# Run on a broker host where kafka-topics.sh is available kafka-topics.sh --bootstrap-server <bootstrap-server> \ --create \ --topic user-profile \ --partitions 3 \ --replication-factor 2 \ --config cleanup.policy=compact \ --config segment.bytes=104857600 \ --config min.cleanable.dirty.ratio=0.5Replace
<bootstrap-server>with your broker address. The topic will now run compaction on each partition.Produce multiple updates for the same key
# Produce three updates for user key "user-123" kafka-console-producer.sh --bootstrap-server <bootstrap-server> \ --topic user-profile \ --property key.serializer=org.apache.kafka.common.serialization.StringSerializer \ --property value.serializer=org.apache.kafka.common.serialization.StringSerializer > user-123 {"name":"Alice"} > user-123 {"name":"Alice A."} > user-123 {"name":"Alice B."}Each line is a new record with the same key but different value. The console producer will exit after you send the records.
Wait for compaction to run
Compaction runs asynchronously. On a busy broker it may take a few minutes. You can monitor the broker logs for messages like
Compaction for user-profile-0 startedor use JMX metricskafka.log:type=Log,name=CompactionRateto see progress.Consume the topic after compaction
# Consume with earliest offset to see the compacted log kafka-console-consumer.sh --bootstrap-server <bootstrap-server> \ --topic user-profile \ --from-beginning \ --property print.key=true \ --property key.deserializer=org.apache.kafka.common.serialization.StringDeserializer \ --property value.deserializer=org.apache.kafka.common.serialization.StringDeserializerYou should see only one record for
user-123– the last value produced. Older records are removed by compaction.
Trade‑offs and Limitations
- Disk Usage During Compaction – Until the compaction job finishes, both old and new segments coexist. This can temporarily double the storage footprint for a partition.
- Delayed Deletion – Compaction is not instant; records may persist until a compaction cycle runs. If you need immediate removal, combine compaction with
retention.msor use tombstones. - CPU & I/O Load – Compaction consumes broker resources. On high‑volume clusters, monitor
CompactionRateandBytesCompactedmetrics to avoid throttling. - Consumer Ordering Guarantees – Compaction preserves record order within a segment but removes older records. Consumers that rely on seeing every historical event will break.
- Key‑Only Records – Records without a key are ignored by compaction. Ensure your producer always sends keys if you expect compaction to apply.
Practical Check: Verify Compaction Effectiveness
After compaction, run:
# List topic configs to confirm cleanup.policy
kafka-configs.sh --bootstrap-server <bootstrap-server> \
--describe \
--entity-type topics \
--entity-name user-profile
Look for cleanup.policy=compact and segment.bytes values. Then inspect the log directory on the broker (e.g., /var/lib/kafka/data/user-profile-0/) for the number of segments. A reduced segment count after compaction indicates successful cleanup.
When to Avoid Compaction
If your application needs a complete audit trail or depends on replaying every event, compaction is unsuitable. In such cases stick to retention.ms or retention.bytes alone. Also, if you have very short‑lived keys that change frequently, compaction may not keep up, leading to stale data.
Actionable Takeaway
Enable cleanup.policy=compact on topics that represent latest‑state snapshots, but pair it with sensible segment.bytes and min.cleanable.dirty.ratio values. Monitor compaction metrics, and if you need to delete a key entirely, emit a tombstone (a record with the key and a null value) and adjust delete.retention.ms accordingly. By doing so you keep your Kafka cluster lean while still providing a reliable source of truth for consumers.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.