Sizing Elasticsearch Data Streams with ILM: An Architecture Note for Log and Metrics Storage
An architecture note on the minimal Elasticsearch data stream + ILM design for logs and metrics: sizing rollover by shard size, trust boundaries, operational checks, and when to redesign.
03 Mar 2026, 19:20 UTC

If you are putting logs or metrics into Elasticsearch, the decision that matters is not which ingest pipeline to use — it is how indices are born, rolled, and deleted. Get that wrong and you end up with thousands of tiny shards, a cluster that spends its life recovering, or a delete phase that destroys data you later need. The supported pattern for append-only time-series data is a data stream managed by Index Lifecycle Management (ILM): writes always land on one backing write index, and rollover creates new backing indices when age, size, or document-count thresholds are hit. This note covers the smallest design that works, where the trust boundaries are, and what would force a redesign. Behavior described here matches Elasticsearch 7.9+ and 8.x; confirm API details against your deployed version, because ILM defaults and data stream APIs differ across majors.
Requirements this design assumes
Data streams with ILM fit when all of the following hold:
- Documents are append-only. You never update or delete individual events by ID (doing so requires targeting backing indices directly, which is a smell).
- Every document has a
@timestampfield; queries are predominantly time-ranged. - Retention is a function of age (delete after N days), not of per-document logic.
- You can tolerate Elasticsearch as a queryable store, not a system of record — more on that below.
If any of these fail, jump to the redesign conditions at the end.
The smallest suitable design
Resist the urge to start with hot/warm/cold/frozen tiers. The minimal production design has exactly three moving parts:
- One ILM policy with a hot phase (rollover) and a delete phase.
- One composable index template matching the data stream name pattern, carrying mappings, shard settings, and the lifecycle policy reference.
- One data stream created from that template.
Example policy, created with PUT _ilm/policy/logs-policy against the cluster (requires manage_ilm privilege):
{
"policy": {
"phases": {
"hot": {
"actions": {
"rollover": {
"max_primary_shard_size": "30gb",
"max_age": "7d"
}
}
},
"delete": {
"min_age": "30d",
"actions": { "delete": {} }
}
}
}
}Note the rollover condition: size or age, whichever comes first. Size your rollover by shard targets, not purely by time. A common working target is primary shards in the tens of gigabytes (roughly 10–50 GB). Oversized shards make recovery and rebalancing slow; undersized shards waste heap and inflate cluster state, which is how clusters die at 50,000 shards. If your daily volume is 5 GB, a daily rollover gives you 5 GB shards — that is a shard-count problem, so roll on size and let indices span multiple days.
The matching index template (PUT _index_template/logs-template, requires manage_index_templates):
{
"index_patterns": ["logs-app-*"],
"data_stream": {},
"template": {
"settings": {
"number_of_shards": 1,
"number_of_replicas": 1,
"index.lifecycle.name": "logs-policy"
},
"mappings": {
"properties": {
"@timestamp": { "type": "date" },
"message": { "type": "text" },
"level": { "type": "keyword" },
"service": { "type": "keyword" }
}
}
}
}Lock mappings down in the template for production. Dynamic mapping on arbitrary log payloads is how you get a mapping explosion — a payload with unbounded keys (user IDs as field names, say) silently creates thousands of fields and eventually rejects writes. Either define explicit fields or set "dynamic": "false" with a catch-all strategy you have thought about.
Trust and data boundaries
Two boundaries deserve explicit treatment:
Elasticsearch is not a system of record by default. The ILM delete phase is irreversible. If there is any chance you need data back — an incident investigation, a compliance request — pair the delete phase with a snapshot policy (SLM) that runs before deletion. Treat snapshots as the archival boundary and the cluster as a query cache over the retention window.
Data streams are not immutable audit storage. Append-only is a convention, not an enforcement mechanism. Anyone with write access to backing indices can modify documents. Compliance workloads need additional controls: restricted index privileges, snapshot-based WORM archival outside the cluster, or a different store entirely.
Operational checks that actually matter
- ILM progress:
GET logs-app-*/_ilm/explain(read-only, safe to run any time) shows which phase and step each backing index is in. This is your first stop when retention is not happening. - Shard allocation:
GET _cluster/healthandGET _cat/shards?v&unassigned. ILM rollover stalls on a yellow/red cluster, so unassigned shards cascade into lifecycle failures. - Disk headroom: watch watermarks. At the flood-stage watermark Elasticsearch flips indices on affected nodes to read-only (
index.blocks.read_only_allow_delete), which stops ingestion into the stream, not just old indices. - Write rejections:
GET _cat/thread_pool/write?v— sustained rejected writes mean saturation, and clients without backoff will lose data. - Rollover sanity: after a rollover, confirm the stream's write index moved:
GET _data_stream/logs-app-prodand check the last backing index.
Failure modes to design around
- ILM stuck at rollover because the cluster is not green, or because something (usually a manual alias operation) pointed the write alias at the wrong index. Diagnose with the explain API; fix allocation first, then retry the step with
POST <index>/_ilm/retry. - Flood-stage read-only block. Recovery requires freeing disk and manually clearing the block (
PUT */_settingswithindex.blocks.read_only_allow_delete: null) — it does not lift itself. - Mapping explosion from dynamic fields, covered above.
- Silent wrong types. A field first indexed as a string stays a string; a later numeric value is rejected or misindexed. Template mappings prevent this.
Verify before you trust it
In staging, create the stream with an aggressive threshold (max_docs: 1000), index sample documents, and confirm a new backing index appears and the write alias moves. Query across the stream before and after rollover to confirm read continuity and that time-range queries only touch expected backing indices. Run the ILM explain API on a backing index and watch it walk the phases. Finally, kill a node or fill a disk in staging and observe allocation, watermark behavior, and recovery time — these are the numbers you will need during the real incident.
Conditions that change the design
- Long-range point-in-time queries over old data → add a frozen tier with searchable snapshots rather than keeping months of data on hot nodes.
- Legal hold or strict retention → move archival to snapshots with independent retention; ILM delete alone is not defensible.
- Frequent updates or high-cardinality entity-centric data → the data stream model no longer fits; use regular indices with aliases and accept the operational overhead of managing rollover yourself.
- Query latency SLAs on recent data with cheap storage for the rest → that is when warm/cold tiers earn their complexity, not before.
Start minimal: one template, one policy, size-based rollover, snapshots before delete. Every tier and phase you add beyond that should trace back to a measured requirement, not a default.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.