Managing State in Kafka: When to Use Log Compaction Over Deletion
Learn how Apache Kafka Log Compaction manages stateful data by retaining the latest value for each key, reducing storage overhead, and implementing tombstones for deletions.
08 Sept 2025, 20:24 UTC

The Storage Dilemma: History vs. State
In a standard Kafka topic, data is treated as a chronological stream. You set a retention period (e.g., 7 days), and once a message hits that age, it is deleted. This works perfectly for event logs or telemetry. However, it fails when you need to store the current state of an object—such as a user's current address or a product's current price.
If you rely on time-based deletion for stateful data, you risk losing the current value of a key simply because it hasn't been updated recently. The solution is Log Compaction. Instead of deleting data based on age, compaction ensures Kafka retains at least the last known value for every single message key.
How the Log Cleaner Operates
Log compaction is handled by a background process called the Log Cleaner. It doesn't scan the entire log constantly; instead, it works on segments (the actual files on disk). When a segment is marked for cleaning, the cleaner reads the records and discards older versions of a key, keeping only the most recent one.
This creates a "materialized view" of your data directly within the log. For developers building Kafka Streams applications or KTables, this is critical. It allows a new instance of a service to bootstrap its local state by reading the compacted topic from the beginning without processing millions of obsolete updates.
Handling Deletions with Tombstones
If the cleaner only keeps the latest value, how do you actually delete a key? You cannot simply stop sending updates, as the last known value would persist forever.
To remove a key, you must produce a Tombstone: a message with the specific key you want to delete and a null value. The Log Cleaner recognizes this null value as a signal to eventually remove all previous entries for that key. Note that the tombstone itself is kept for a configurable period (delete.cleanup.min.interval.ms) to ensure that downstream consumers have time to see the deletion event before the record vanishes entirely.
Practical Implementation: Configuring a Compacted Topic
To implement this, you must change the cleanup.policy from the default delete to compact. You can also use both (compact,delete) if you want to keep the latest value but still purge data that hasn't been updated in years.
Run the following command on your broker or a client with administrative permissions to create a state-tracking topic:
# Create a topic specifically for state management
bin/kafka-topics.sh --bootstrap-server localhost:9092 --create \
--topic user-profiles \
--partitions 3 \
--replication-factor 1 \
--config cleanup.policy=compact
Verification Steps:
- Produce three messages with the same key
user123but different values (e.g., "New York", "London", "Tokyo"). - Wait for the Log Cleaner cycle to trigger (this is not instantaneous).
- Consume from the beginning:
bin/kafka-console-consumer.sh --bootstrap-server localhost:9092 --topic user-profiles --from-beginning. - You should see only the final value ("Tokyo") remaining for
user123.
Trade-offs and Performance Risks
Log compaction is not a "free" feature. It introduces specific overheads that can impact broker health if not monitored:
- Disk I/O Spikes: The cleaning process involves reading and rewriting segments. On a heavily loaded broker, this can compete with produce/consume requests for disk throughput.
- CPU Load: The Log Cleaner requires CPU cycles to track offsets and manage the cleaning queue.
- Lag: Compaction is asynchronous. There is a window of time where old values still exist in the log before the cleaner reaches that segment. Your application logic must be able to handle seeing multiple versions of a key during a replay.
Decision Summary
Choose Log Compaction when your primary goal is to maintain a current snapshot of keys (State) rather than a history of changes (Events). If you need both, consider using two separate topics: one for the raw event stream (cleanup.policy=delete) and one for the compacted state (cleanup.policy=compact), updating the latter via a stream processor.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.