Using Logstash Dead Letter Queue to Capture and Reprocess Failed Events
Learn how to enable Logstash's Dead Letter Queue to persist failed events, inspect them, and safely reprocess without losing data.
30 Aug 2026, 19:03 UTC

Problem: a single bad line can silence your pipeline
In a production Logstash pipeline, a malformed JSON entry or a filter timeout can cause an event to be dropped or raise an exception. When that happens, downstream systems lose visibility into the offending record and the pipeline may stall, leading to data loss or delayed alerts.
Thesis: the Dead Letter Queue (DLQ) preserves every unprocessable event for later inspection
Enabling the DLQ tells Logstash to write the original event, plus a small amount of metadata, to a configurable storage location whenever a pipeline encounter a failure. The main pipeline continues with the next event, so throughput is not interrupted while you retain a durable copy of the failure for replay.
How the DLQ works
When a filter, codec, or output returns an error, Logstash:
- Serializes the original event as a JSON line.
- Adds metadata fields such as
[_metadata][dead_letter_queue][entry_time]and[_metadata][dead_letter_queue][reason]. - Appends the combined record to a file (or Redis list) under the configured DLQ path.
- Returns control to the pipeline to process the next incoming event.
The DLQ is therefore a simple append‑only log; ordering of replayed events relative to the live stream is not guaranteed, which downstream consumers must handle.
Configuration steps
Edit logstash.yml (typically found in $LOGSTASH_HOME/config/) and add or adjust the following lines:
# Enable the DLQ feature
dead_letter_queue.enable => true
# Choose a storage backend; here we use the local filesystem
dead_letter_queue.max_bytes => 1024mb
# Optional: set a custom path if the default $LOGSTASH_DATA/dead_letter_queue is unsuitable
dead_letter_queue.path => "${LS_DATA_DIR}/dead_letter_queue"
After saving the file, restart Logstash. The service needs read/write permission to the directory pointed at by dead_letter_queue.path. Verify the restart succeeded with:
# Run as the user that owns the Logstash installation
sudo systemctl restart logstash
# Check that the DLQ directory exists
ls -ld $LOGSTASH_DATA/dead_letter_queue
You should see a directory owned by the Logstash user with default permissions drwxr-xr-x.
Worked example: inject a broken JSON line and replay it
1. Generate a malformed event using the stdin input (run this in a terminal with the same user that runs Logstash):
echo '{ "msg": "test"' | bin/logstash -e 'input { stdin { } } output { stdout { codec => rubydebug } }'
Because the JSON is missing a closing brace, the stdin codec will fail to parse the line. Logstash will:
- Skip emitting the event to the stdout output.
- Write the original line plus DLQ metadata to a new file under
$LOGSTASH_DATA/dead_letter_queue/main.log.
2. Confirm the DLQ entry (no invented output; just the check you would perform):
# List the DLQ files; you should see a newly created segment
ls -l $LOGSTASH_DATA/dead_letter_queue/
# Inspect the first line (you will see the original JSON and added metadata)
head -n 1 $LOGSTASH_DATA/dead_letter_queue/main.log
3. Replay the event with a second pipeline that reads from the DLQ input, attempts to fix the JSON, and sends the result to stdout:
# Create a temporary config file replay.conf
cat > replay.conf <<'EOF'
input {
dead_letter_queue {
path => "${LS_DATA_DIR}/dead_letter_queue"
commit_offsets => false
sincedb_path => "/dev/null" # start from the beginning each run
}
}
filter {
# Attempt to repair the JSON; if it still fails, the event will go to the DLQ again
mutate {
gsub => [ "message", "\\'$\", "\"" ]
}
# Example: add a missing closing brace if the line ends with a colon or quote
if [message] !~ /\}$/ {
mutate { add => { "message" => "%{message}" } }
}
}
output {
stdout { codec => rubydebug }
}
EOF
# Run the replay pipeline
bin/logstash -f replay.conf
When the replay pipeline reads the DLQ entry, it will apply the mutate filter and output a corrected event to the terminal. This demonstrates end‑to‑end reprocessing without interrupting the original pipeline.
Trade‑off and limitation
The DLQ consumes disk space roughly proportional to the volume of failed events. A burst of malformed data can quickly fill the partition hosting dead_letter_queue.path, causing the pipeline to pause when the storage becomes unavailable (back‑pressure). To mitigate this:
- Monitor the DLQ size with a simple metric:
du -sb $LOGSTASH_DATA/dead_letter_queue. - Set a retention policy (e.g., delete files older than 7 days) using a cron job or Logstash’s
pipeline.batch.sizeandpipeline.workerstuning to keep the queue bounded. - Alert on rapid growth; a threshold of 80 % of the allocated
dead_letter_queue.max_bytesis a practical starting point.
Additionally, replayed events may arrive out of order or as duplicates relative to the live stream; downstream systems should be made idempotent or equipped with deduplication logic.
Actionable closing
Turn on the DLQ for every critical Logstash pipeline, verify the directory appears after restart, and inject a known bad line to confirm capture. Set up routine checks on queue size and schedule a nightly replay job that reads from the DLQ, applies corrective filters, and writes to a stable output. By treating the DLQ as a safety net rather than a debugging curiosity, you gain visibility into failure modes and the ability to recover lost data without sacrificing pipeline availability.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.