Choosing Between Logstash and Elasticsearch Outputs for Beats
A decision guide for selecting the Beats output plugin based on data volume, processing needs, and latency requirements.
20 Mar 2026, 17:55 UTC

Decision: Where should Beats send its data?
When designing a Beats‑based ingestion pipeline you must choose an output plugin. The two primary options are sending events directly to an Elasticsearch cluster or routing them through a Logstash instance first. The choice hinges on data volume, required transformation complexity, latency tolerance, and operational overhead.
Constraints to consider
- Data volume: High‑throughput streams can overwhelm Elasticsearch if there is no buffering layer.
- Processing needs: Simple field mapping or no transformation works well with the direct output; complex grok, conditionals, or external lookups require Logstash.
- Latency: Direct output adds minimal hop latency; Logstash introduces processing delay.
- Operational footprint: Direct output reduces the number of moving parts; Logstash adds a separate service to monitor and scale.
Comparison of supported options
| Feature | Direct Elasticsearch output | Logstash output |
|---|---|---|
| Infrastructure overhead | Only Beats and Elasticsearch | Beats, Logstash, Elasticsearch |
| Transformation capability | Basic fields, drop/rename, simple processors | Full Logstash plugin set (grok, mutate, ruby, jdbc, etc.) |
| Latency (typical) | Low (network round‑trip only) | Higher (queue + filter pipeline) |
| Buffering during spikes | None (back‑pressure applied to Beats) | Logstash persistent queues can absorb bursts |
| Security | SSL/TLS, basic auth, API key, PKI | Same options for Beats→Logstash and Logstash→Elasticsearch links |
Trade‑off summary
- Choose direct Elasticsearch output when:
- Events are already well‑structured (e.g., JSON logs).
- Processing limited to field dropping, renaming, or simple ingest pipelines.
- Volume is moderate or you can rely on Elasticsearch’s bulk API back‑pressure.
- Minimizing latency and operational complexity is a priority.
- Choose Logstash output when:
- You need to combine data from multiple Beats sources.
- Enrichment via external databases, lookup tables, or complex conditional logic is required.
- Traffic spikes are common and you want a buffering layer to protect Elasticsearch.
- You already operate a Logstash cluster for other pipelines.
Concrete implementation examples
The following snippets show a minimal Filebeat configuration for each output. Replace placeholders with your actual values.
Direct Elasticsearch output
# filebeat.yml
filebeat.inputs:
- type: filestream
enabled: true
paths:
- /var/log/myapp/*.log
output.elasticsearch:
hosts: ["${ES_HOST}:9200"]
username: "${ES_USER}"
password: "${ES_PASSWORD}"
# Optional: enable SSL/TLS verification
ssl.certificate_authorities: ["/path/to/ca.crt"]
index: "filebeat-%{+yyyy.MM.dd}"
After starting Filebeat (sudo systemctl start filebeat), you can verify that documents appear in Elasticsearch:
# Run on any host with curl access to the cluster
curl -s -u "${ES_USER}:${ES_PASSWORD}" "${ES_HOST}:9200/filebeat-*/_search?size=1&pretty" | jq .hits.total.value
If the configuration is correct, the response will show a non‑zero total (the exact number depends on your log volume). No specific numeric value is claimed here; you should observe that the count reflects the lines Filebeat has read.
Logstash output
# filebeat.yml
filebeat.inputs:
- type: filestream
enabled: true
paths:
- /var/log/myapp/*.log
output.logstash:
hosts: ["${LS_HOST}:5044"]
# Optional TLS
ssl.certificate_authorities: ["/path/to/ls-ca.crt"]
On the Logstash side, a simple pipeline (/etc/logstash/conf.d/beats.conf) might look like:
input {
beats {
port => 5044
ssl => true
ssl_certificate => "/path/to/ls.crt"
ssl_key => "/path/to/ls.key"
}
}
filter {
# Example: parse a common Apache log format
grok {
match => { "message" => "%{COMBINEDAPACHELOG}" }
}
# Example: enrich with a static field
mutate { add_field => { "env" => "production" } }
}
output {
elasticsearch {
hosts => ["${ES_HOST}:9200"]
user => "${ES_USER}"
password => "${ES_PASSWORD}"
index => "filebeat-%{+yyyy.MM.dd}"
}
}
Start Logstash (sudo systemctl start logstash) and then verify the enriched fields:
curl -s -u "${ES_USER}:${ES_PASSWORD}" "${ES_HOST}:9200/filebeat-*/_search?size=1&pretty" | jq '.hits.hits[0]._source | {message, env}'
You should see the env field present with the value you set in the filter, indicating that Logstash performed the transformation.
Limitations and practical checks
- Direct output risk: Sustained high traffic can cause Elasticsearch to reject bulk requests, leading to back‑pressure that may stall Filebeat. Monitor the
output.elasticsearch.eventsandoutput.elasticsearch.failuresmetrics via Filebeat’s internal monitoring or Metricbeat. - Logstash lag: Even with persistent queues, each event incurs filter processing time. Measure end‑to‑end latency by adding a timestamp field at the source and comparing it to the
@timestampfield after indexing. - Verification tip: Use the Elasticsearch
_cat/indices?vAPI to confirm index growth rate matches expected event volume. A sudden drop in growth may indicate output problems. - Rollback: Changing the output plugin only requires editing the Beats configuration and restarting the service. Reverting to the previous output is as simple as restoring the earlier
filebeat.ymland restarting.
Summary
If your pipeline needs only light field adjustments and you can tolerate direct back‑pressure on Elasticsearch, the Elasticsearch output offers the lowest latency and simplest ops. When you require complex enrichment, multi‑source aggregation, or a buffer to absorb traffic spikes, route Beats through Logstash despite the added latency and operational overhead. Validate your choice by checking index existence, document counts, and, for Logstash, the presence of transformed fields after a short test run.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.