Filebeat and Elasticsearch: Event Duplication Risks During Network Retries
0 reputation · 04 Aug 2025, 22:33 UTC
0 reputation · 04 Aug 2025, 22:33 UTC
Filebeat utilizes an at-least-once delivery model to ensure data integrity. By default, the agent waits for an acknowledgment from Elasticsearch before removing events from its internal queue. When a bulk request fails, the agent employs exponential backoff and splits the bulk to retry only the failed items.
There is a design consideration regarding the pipeline.ack feature. Shifting the acknowledgment responsibility to the Beats pipeline can reduce latency but may increase the risk of duplicate indexing if the agent shuts down abruptly before the pipeline clears.
pipeline.ack is enabled, as events remain in memory longer.flush.min_events and network failures may lead to events being flushed before the acknowledgment is received.Given these behaviors in Beats 7.x and later, how does the system resolve the conflict between minimizing latency via pipeline acknowledgments and preventing duplicate writes during a crash? Under what specific conditions does the bulk_max_size configuration influence the likelihood of duplicates upon restart?
29275 reputation · 05 Aug 2025, 01:52 UTC
Filebeat avoids duplicate writes after a crash by relying on its persistent registry (file.registry) that records the last successfully acknowledged offset for each input; on restart it resends only events whose offsets are not marked as acknowledged.
pipeline.ack is enabled, acknowledgment is moved to the internal Beats pipeline, reducing network round‑trip latency but keeping events in memory until the pipeline flushes.bulk_max_size defines the maximum number of events in a single Elasticsearch bulk request; larger values increase the number of events that can be in flight at any moment.If pipeline.ack is enabled, Filebeat considers an event sent once the pipeline has processed it, which lowers latency. However, the event remains in memory until the pipeline flushes and the registry is updated. A sudden crash before the pipeline flush or before the registry write can leave the event unacknowledged on disk, causing Filebeat to resend it after restart and potentially create a duplicate. The persistent registry mitigates this because, upon restart, Filebeat checks the stored offset; only events whose offsets are not recorded as acknowledged are resent. Thus duplicates occur only when the registry itself is lost, corrupted, or not yet updated for the events in flight.
bulk_max_size influences duplicate likelihoodA larger bulk_max_size means more events are held in a single bulk request. When a network interruption occurs after the bulk is sent but before Elasticsearch replies, the entire bulk may be retried. If pipeline.ack is enabled and the process crashes before the pipeline acknowledges those events, the larger batch increases the number of events that could be resent, raising duplicate risk. Conversely, a smaller bulk_max_size limits the batch size, reducing the window of potential duplication.
pipeline.ack is enabled in your filebeat.yml.file.registry location and ensure it resides on reliable storage (not a volatile tmpfs).bulk_max_size (e.g., from 5000 to 1000) and monitor heap usage.-debug) and simulate a network drop to observe pipeline acknowledgment and registry updates.To give a definitive recommendation on whether adjusting bulk_max_size will lower duplicate risk, we need to know if pipeline.ack is currently enabled. If it is disabled, the acknowledgment already waits for Elasticsearch, and bulk size influences duplicate risk only through network retries, not pipeline latency.
Use comments to ask for clarification. Post a solution as an answer.
29,275 reputation · 05 Aug 2025, 08:48 UTC
To resolve the conflict between latency and duplication, it is important to distinguish between at-least-once delivery and idempotent indexing. While pipeline.ack and bulk_max_size influence when a retry is triggered, they do not prevent duplicates if the retry occurs after Elasticsearch has already committed the data but before the acknowledgment reaches Filebeat.
In Beats 7.x+, the most effective way to ensure idempotency during network retries or crashes is by configuring a deterministic document_id. Without a specified ID, Elasticsearch generates a unique identifier for every request, treating every retry as a new event.
document_id setting in the Elasticsearch output to derive an ID from a unique field in the log (e.g., a UUID or a combination of hostname and timestamp).iptables; with a deterministic ID, the document count in Elasticsearch should remain constant despite multiple retries.