Guide
Diagnosing Logstash Dead Letter Queue Growth and Event Loss
Learn how to spot DLQ‑related event loss, diagnose the root cause, apply fixes, and know when to escalate.
Published by Tasadduq Burney
12 Apr 2026, 17:11 UTC
3 min28.8K views0

Recognizable Condition
Events disappear from the pipeline outputs while Logstash logs repeatedly show messages such as "DLQ enabled", "Failed to send event to output", and "Event routed to DLQ". At the same time, the observed throughput drops and the directory configured for the dead‑letter‑queue (DLQ) grows in size over time.
Cause/Diagnostic Table
| Observed Symptom | Likely Cause | Log / Metric Indicator |
|---|---|---|
| No events reach configured outputs | Output plugin repeatedly fails after exhausting retries | Log lines: "Failed to send event to output" followed by "Event routed to DLQ" |
| Steady increase in DLQ file size | Events are being written to DLQ faster than they can be replayed or removed | Growing size reported by du -sh on the DLQ path or rising dead_letter_queue.bytes in the monitoring API |
| Pipeline stalls or crashes when DLQ is disabled | Unrecoverable output errors cause Logstash to halt processing | Logstash exits with non‑zero status or shows "Pipeline worker killed" when dead_letter_queue.enable: false |
Ordered Checks
- Search the Logstash log for DLQ activity:
grep -i 'dead_letter_queue\|DLQ\|routed to DLQ' /var/log/logstash/logstash.log - Confirm the DLQ configuration in
logstash.ymlor pipeline settings: look fordead_letter_queue.enabled: trueand note thepath.dead_letter_queuevalue. - List the DLQ directory to verify new segment creation:
ls -lt $(grep path.dead_letter_queue logstash.yml | cut -d'=' -f2)– expect recent timestamps and growing file count. - Measure DLQ size over time: run
du -sh $(grep path.dead_letter_queue logstash.yml | cut -d'=' -f2)at intervals or query the monitoring API endpoint/_node/stats/pipeline?prettyfordead_letter_queue.bytes. - Inspect a sample DLQ event to see the original error:
bin/logstash dlq --path --show --limit 1(replacewith the actual directory). The output will include the original exception or error message. - After hypothesizing a fix (e.g., correcting output credentials or network reachability), replay a DLQ event to test:
bin/logstash dlq --path --input stdin --output stdoutand pipe a single event through, observing whether it now succeeds.
Fixes Tied to Findings
- If logs show repeated "Failed to send event to output" with authentication or connection errors, verify the output plugin’s connection details (host, port, credentials, TLS) and correct them in the pipeline configuration.
- If the DLQ size grows because the output is temporarily unavailable (e.g., downstream service downtime), increase the output’s
retry_countorretry_intervalsettings, or implement a buffering queue (e.g., Redis) before the Logstash output. - If the DLQ segment files are numerous and old, enable automatic DLQ cleanup by setting
dead_letter_queue.max_bytesor periodically runbin/logstash dlq --path --cleanup --older-than 7dto remove stale segments. - When the output plugin itself is misconfigured (wrong codec, missing required fields), adjust the pipeline to match the output’s expectations, then replay the DLQ to confirm events now pass through.
- DLQ disk usage exceeds 80 % of the volume holding
path.dead_letter_queueand continues to rise despite corrective actions. - Replaying DLQ events still results in the same error after verifying output connectivity, plugin versions, and pipeline syntax.
- Logstash process repeatedly crashes or is killed by the OOM killer while DLQ is disabled, indicating a non‑recoverable fault that requires deeper plugin or JVM investigation.
- Monitoring shows sustained pipeline latency > 5 s and throughput < 10 % of baseline for more than 15 minutes after attempted fixes.
Escalation Criteria
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.