Stop Deleting Log Indices by Hand: A Practical Case for Elasticsearch ILM
Manual index cleanup scripts are a liability. Elasticsearch ILM automates rollover, phase transitions, and retention for time-series data — here's a practical rollover example and the trade-offs to know first.
30 Jun 2026, 11:52 UTC

If you run Elasticsearch for logs or metrics, you eventually hit the same wall: indices pile up, shard counts creep past what your heap can comfortably hold, and somebody ends up writing a cron job that deletes old indices by name pattern. It works until it doesn't — the cron deletes the wrong index, or retention quietly changes and nobody updates the script.
Index Lifecycle Management (ILM) is Elasticsearch's built-in answer to this. It's a supported feature that automates rollover, phase transitions, and deletion for time-series data, declared as a policy rather than scripted as cleanup. This post makes the case for adopting it incrementally, with a concrete rollover example you can adapt.
The problem ILM actually solves
Time-series indices have a natural lifecycle: they're written heavily at first, queried often for a few days, then mostly sit around until they expire. Managing that manually means tracking three things yourself — when to stop writing to an index and start a new one (rollover), when to move or optimize older data, and when to delete it. Each is easy to get subtly wrong. Roll over too aggressively and you drown the cluster in tiny shards; too loosely and you end up with oversized shards that are slow to search and painful to recover.
ILM moves those decisions into a policy attached to the index template, so every new index inherits the same behavior. The cluster — not your cron — decides when an index rolls over, cools down, or disappears.
Phases in plain language
An ILM policy is organized into phases:
- Hot — the index is actively written. This is where rollover happens, triggered by age, size, or document count.
- Warm — no longer written, but still queried. Typical actions: force merge to fewer segments, shrink to fewer shards, or allocate to cheaper nodes.
- Cold / frozen — rarely queried, kept for compliance or occasional lookups, usually on slower storage.
- Delete — the index is removed after a retention period.
The key point: you don't need every phase. Many deployments are well served by hot plus delete alone. Adding warm and cold phases only pays off when you actually have tiered hardware or distinct query patterns.
A worked example: application logs, 30-day retention
Goal: daily app logs roll over when the write index reaches roughly 1 GB or 7 days, and are deleted 30 days after rollover. Run these against your cluster with a user that has manage_ilm and index management privileges (e.g., via Kibana Dev Tools or curl).
First, define the policy:
PUT _ilm/policy/logs-30d
{
"policy": {
"phases": {
"hot": {
"actions": {
"rollover": {
"max_primary_shard_size": "1gb",
"max_age": "7d"
}
}
},
"delete": {
"min_age": "30d",
"actions": { "delete": {} }
}
}
}
}Then an index template that applies the policy and sets the rollover alias:
PUT _index_template/logs-template
{
"index_patterns": ["logs-*"],
"template": {
"settings": {
"index.lifecycle.name": "logs-30d",
"index.lifecycle.rollover_alias": "logs-write"
}
}
}Finally, bootstrap the first index with the write alias — this step is mandatory, and getting it wrong is the most common ILM failure:
PUT logs-000001
{
"aliases": {
"logs-write": { "is_write_index": true }
}
}All indexing goes to logs-write. When the active index hits either condition, ILM creates logs-000002, repoints the alias, and starts the 30-day clock on the old index. Note that max_primary_shard_size was introduced in 7.13; on older versions use max_size or max_docs instead — confirm which rollover conditions your deployed version supports before copying this.
The trade-offs
ILM is not free of sharp edges:
- Rollover sizing is a real decision. A 1 GB threshold on a low-volume stream creates many small shards; a 50 GB threshold on a busy one creates oversized ones. Aim for shards in the tens of GB for search-heavy workloads, and adjust after observing actual volume.
- Delete is destructive. If retention data must be recoverable, pair ILM with Snapshot Lifecycle Management so a snapshot exists before the delete phase fires.
- Policy edits aren't fully retroactive. Indices already attached to a policy keep their existing phase definitions in some respects; changing a policy doesn't magically re-phase old indices. Plan migrations explicitly rather than assuming the edit propagates.
- Rollover depends on the write alias. Missing or misconfigured
is_write_indexmeans indexing or rollover fails — verify the bootstrap step before trusting the automation.
Start small and verify
The pragmatic rollout: begin with hot (rollover) plus delete only. Watch the cluster for a couple of retention cycles, then consider warm-phase actions like force merge once you understand query and storage patterns.
To confirm it's working, ask the cluster rather than guessing:
GET logs-000001/_ilm/explainThis shows the current phase, action, and step for the index. Also check GET logs-000001/_settings for the policy name and rollover alias, and compare shard count and index age before and after a rollover. If you want to see a transition quickly in a test cluster, create a throwaway policy with a tiny max_docs value, index documents past the threshold, and wait for the ILM poll interval (default 10 minutes, tunable via indices.lifecycle.poll_interval) — then delete the test policy and indices afterward.
The payoff for an afternoon of setup: retention becomes a declared, reviewable policy instead of a script someone wrote two years ago and nobody wants to touch.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.