Keeping Logs Safe: Using Logstash Persistent Queues for At‑Least‑Once Delivery
Learn how Logstash persistent queues protect in‑flight events during restarts, with a concrete configuration example, verification steps, and trade‑offs to consider.
07 Sept 2025, 04:27 UTC

The problem: lost logs during a restart
When a Logstash pipeline stops—whether because of a planned upgrade, a JVM crash, or a node drain—any events that are sitting in memory between the input and filter stages disappear. Downstream systems then see gaps, and dashboards show missing data. This is especially painful for bursty workloads where a short interruption can wipe out thousands of log lines.
Thesis: persistent queues give you a safety net
Enabling Logstash’s persistent queue writes each incoming event to disk before it reaches the filter stage. If the process stops, the queue survives on disk and is replayed when Logstash starts again, guaranteeing at‑least‑once delivery for every event that made it into the queue.
How to enable a persistent queue
Add a few settings to your pipeline configuration file (e.g., logstash.conf). The queue type must be persisted, and you should size the queue based on your expected event size and throughput.
# logstash.conf
input {
tcp {
port => 5000
codec => json
}
}
filter {
# your filters here
}
output {
elasticsearch {
hosts => ["http://es:9200"]
index => "logs-%{+YYYY.MM.dd}"
}
}
# Queue settings (placed outside any plugin block)
queue.type: persisted
queue.page_capacity: 64mb
queue.max_bytes: 1gb
path.queue: /var/lib/logstash/queue
Place the configuration in /etc/logstash/conf.d/ (or wherever your Logstash instance reads configs) and restart the service.
Worked example: protecting a burst of TCP events
- Start Logstash with the config above.
- Using a tool like
ncor a simple script, send a burst of 10 000 JSON events tolocalhost:5000. - While the burst is still being ingested, stop Logstash (
systemctl stop logstashor kill the JVM). - Check the queue directory (
path.queue)—you should see numbered.logfiles containing the events that were accepted but not yet processed. - Restart Logstash.
- Observe the output (e.g., Elasticsearch) and verify that the same 10 000 events appear after the restart.
You can confirm the replay via the monitoring API: GET _node/stats/pipeline?pretty shows queue.events dropping to zero once the queued events have been flushed.
Trade‑off: disk I/O and storage planning
Persistent queues add latency because each event must be written to disk before filtering. On a fast SSD this overhead is usually modest (a few milliseconds), but on network‑mounted or slow disks it can become a bottleneck and cause the input thread to block. If the queue fills beyond queue.max_bytes, Logstash will back‑pressure the upstream source, potentially leading to dropped events at the producer side.
To mitigate these risks:
- Measure average event size and peak events‑per‑second.
- Allocate a dedicated, low‑latency SSD for
path.queue. - Set
queue.max_bytesto a value that leaves ample free space (e.g., 60‑70 % of the disk). - Monitor the metrics
queue.eventsandqueue.memory_eventsvia the Logstash monitoring API or Beats/Metricbeat dashboards.
Actionable closing
Before turning on persistent queues in production, run a short load test that mimics your real traffic pattern. Verify that the queue drains correctly after a simulated stop and that end‑to‑end latency stays within your SLA. Once you have confidence, enable the feature, allocate fast local storage, and keep an eye on the queue depth metrics. This simple change can turn an unpredictable log loss scenario into a reliable, at‑least‑once pipeline.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.