Choosing a Multi‑Line Log Strategy for Filebeat 8.x
Compare Filebeat’s built‑in multiline codec, the new filestream input, processor‑based parsing, and sidecar shippers to pick the right approach for Java stack traces and container workloads.
12 Jul 2025, 10:30 UTC

Decision and constraints
Teams that ship Java stack traces, multi‑line application logs, or container stdout/stderr to Elasticsearch must decide where the line‑merging happens. The choice is driven by three constraints:
- Input type – legacy
loginput vs. GAfilestreaminput (8.10+). - Deployment model – static
filebeat.yml, Fleet‑managed Elastic Agent, or Kubernetes autodiscover. - Operational budget – CPU on the application node, extra hop latency, and config‑management overhead.
If you control the Filebeat configuration directly and run the log input, the built‑in multiline codec is the simplest, lowest‑latency path. If you have migrated to filestream or use Fleet, the syntax changes. When you cannot touch the agent (e.g., regulated hosts) or need to offload merging from high‑throughput nodes, a sidecar shipper such as Fluent Bit or Vector becomes attractive.
Supported options at a glance
| Option | Where merging occurs | Config surface | Best fit | Key limitation |
|---|---|---|---|---|
| Filebeat multiline codec (log input) | Harvester, before any processor | filebeat.yml → inputs.multiline.* | Static hosts, Java stack traces, low‑to‑moderate EPS | Not compatible with filestream input; regex backtracking risk |
| Filebeat filestream multiline | Filestream harvester, separate multiline section | filebeat.yml → inputs.filestream.multiline.* | New deployments on 8.10+, same host constraints as above | Different pattern syntax; migration required |
| Processors (decode_json_fields after multiline) | After merging, inside the pipeline | processors.decode_json_fields | JSON payloads that become valid only after merging | Cannot fix fragmented JSON without prior merge |
| Kubernetes autodiscover hints | Per‑pod, at harvester start | Pod annotations co.elastic.logs/multiline.* | Dynamic container workloads, zero‑touch config | Requires Elastic Agent/Filebeat with autodiscover enabled |
| Sidecar shipper (Fluent Bit / Vector) | In a separate container, before forwarding | Sidecar config (e.g., Fluent Bit multiline_parser) | High‑throughput nodes, regulated hosts, polyglot pipelines | Extra hop, more components to operate |
Trade‑offs
- Performance – The codec buffers lines in memory until the pattern matches or
max_lines/timeouttriggers. Large stack traces (thousands of lines) can pressure the harvester; setmax_lines: 500andtimeout: 5sconservatively. - Regex safety – Patterns are evaluated per line. Avoid look‑ahead/look‑behind; prefer anchored expressions such as
^\d{4}-\d{2}-\d{2}. - Filter interaction –
exclude_lines/include_linesrun after merging. A pattern that matches only the first line will not exclude the merged event if later lines contain the excluded term. - Operational complexity – Sidecars add a deployment unit and a network hop but keep the application node CPU‑light. Fleet‑managed Elastic Agent centralises config but uses a different schema (
streams.multiline), so legacyfilebeat.ymlis ignored. - Throughput ceiling – Above ~10k EPS the harvester can become a bottleneck. In that regime, ship raw lines to Kafka and merge in Logstash or an Elasticsearch ingest pipeline.
Concrete implementation: Filebeat multiline codec for Java stack traces
The following snippet shows a minimal filebeat.yml for a host that writes Java logs to /var/log/app/*.log. It merges lines that start with a timestamp (ISO‑8601) and treats everything until the next timestamp as a single event.
filebeat.inputs:
- type: log
enabled: true
paths:
- /var/log/app/*.log
multiline.pattern: '^\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}'
multiline.negate: true
multiline.match: after
multiline.max_lines: 500
multiline.timeout: 5s
# Optional: drop noisy debug lines after merging
# exclude_lines: ['^\s*DEBUG']
output.elasticsearch:
hosts: ["https://es.example.com:9200"]
api_key: "${ES_API_KEY}"
Run the configuration test before deploying:
# Run as the filebeat user (or root if you manage the service)
filebeat test config -c /etc/filebeat/filebeat.yml
Expected result: Config OK with no regex compilation errors.
Validation steps
- Create a test log file containing a known multi‑line event (timestamp line + 30‑line stack trace).
- Run Filebeat once with console output to count emitted events:
The count should equal the number of logical log entries (e.g., 1 for the test file).filebeat -c /etc/filebeat/filebeat.yml -once -e \ -d "*" 2>&1 | grep -c '"message"' - Inspect the merged JSON using
output.console.pretty: truein a temporary config and verify that themessagefield contains the full stack trace without line breaks. - Check harvester metrics via the monitoring endpoint (
http://localhost:5066/stats) forfilebeat.harvester.open_filesandlibbeat.output.events.ackedto ensure no back‑pressure builds up.
Limitations and when to reconsider
- If you have migrated to the
filestreaminput, rewrite the multiline block underinputs.filestream.multiline– the pattern syntax is identical but the nesting differs. - For Kubernetes clusters, prefer pod annotations (
co.elastic.logs/multiline.pattern,co.elastic.logs/multiline.negate: "true") so each team owns its pattern without a central config rollout. - When EPS exceeds the harvester’s comfortable range, move merging to a downstream Logstash pipeline or an Elasticsearch ingest processor; this also decouples schema changes from the agent.
- Regulated environments that forbid additional daemons on the application host may still run a sidecar in the same pod (Kubernetes) or a privileged container on the host; evaluate the security policy before adding Fluent Bit or Vector.
By matching the merging strategy to your input type, deployment model, and throughput profile you avoid the common pitfalls of over‑buffering, regex catastrophes, and filter‑order surprises.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.