Managing State in Kafka: When to Use Log Compaction Over Retention
Learn how to use Apache Kafka Log Compaction to manage stateful data, handle deletions with tombstones, and avoid the pitfalls of infinite log growth.
20 Jun 2026, 13:58 UTC

The Problem: The Infinite Growth of State
\nWhen using Apache Kafka to store the current state of an object—such as a user profile, a device configuration, or an account balance—standard time-based or size-based retention policies fail. If you set a 7-day retention period, you lose the state of any object that hasn't been updated in a week. If you set retention to infinite, your disk will eventually fill up with thousands of obsolete updates for the same keys.
\nThe solution is Log Compaction. Instead of deleting data based on age, Kafka retains the last known value for every unique key. This transforms a Kafka topic from a simple stream of events into a durable, distributed key-value store.
\nHow Log Compaction Works
\nLog compaction is managed by a background process called the Log Cleaner. Unlike standard deletion, which chops off the end of a log segment, the cleaner scans segments and removes records that have been superseded by a newer record with the same key.
This process is not instantaneous. Kafka doesn't delete the old record the moment a new one arrives; instead, it waits until a specific percentage of the log is \"dirty\" (containing uncompacted records). This is controlled by the min.cleanable.dirty.ratio setting. If this is set to 0.1, the cleaner triggers once 10% of the log consists of records that could potentially be compacted.
Handling Deletions with Tombstones
\nSince compaction only keeps the last value, you cannot simply stop sending updates to delete a key. To remove a key entirely, you must produce a Tombstone: a record with the target key and a null value. The Log Cleaner sees the null value and knows that this key should be removed from the log entirely during the next compaction cycle.
Implementation Example: Creating a State Topic
\nTo implement a compacted topic, you must set the cleanup.policy to compact. You can do this during topic creation or update an existing topic via the CLI.
Run this command on a machine with Kafka binaries installed:
\n# Create a compacted topic for user profiles\nbin/kafka-topics.sh --bootstrap-server localhost:9092 --create --topic user-profiles --partitions 3 --replication-factor 2 --config cleanup.policy=compact\n\nVerifying Compaction Logic
\nTo verify that compaction is working, follow this sequence of operations:
\n- \n
- Produce updates: Send three messages with the key
user_123: first\"Active\", then\"Away\", then\"Offline\". \n - Produce a tombstone: Send a message with key
user_123and anullvalue. \n - Wait for the cleaner: The Log Cleaner runs periodically. You can force it by reducing
min.cleanable.dirty.ratioto a very low value. \n - Consume from beginning: Run the console consumer to see what remains. \n
# Consume from the start to see the final state\nbin/kafka-console-consumer.sh --bootstrap-server localhost:9092 --topic user-profiles --from-beginning --property print.key=true\n\nExpected Result: After compaction, you should see only the final value (or nothing, if the tombstone has been fully processed), rather than the full history of state changes.
\nCritical Trade-offs and Limitations
\n| Factor | \nBehavior | \nRisk | \n
|---|---|---|
| Latency | \nAsynchronous cleanup | \nDisk usage spikes between cleaning cycles. | \n
| Key Uniqueness | \nKey-based retention | \nIf you use non-unique keys, you will lose all historical data except the last record. | \n
| Tombstone Life | \ndelete.retention.ms | \nIf tombstones are deleted too quickly, consumers that are lagging may never see the deletion event. | \n
One significant limitation is that compaction does not guarantee the immediate removal of data. If your application requires a strict \"hard delete\" for compliance (like GDPR), relying solely on the Log Cleaner is insufficient because the old data persists in the log until the cleaner reaches that segment.
\nClosing Summary
\nLog compaction is the primary mechanism for implementing KTable patterns where the topic acts as a materialized view of a database table. To succeed with this pattern, ensure your keys are truly unique, monitor your delete.retention.ms to prevent consumers from missing tombstones, and remember that disk space is reclaimed in cycles, not in real-time.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.