Automating Index Cleanup with Elasticsearch ILM Rollover Policies
Learn how Elasticsearch ILM automates index rollover and deletion based on size or age, reducing manual cleanup and preventing disk‑related cluster issues.
21 Feb 2026, 22:06 UTC

Problem: uncontrolled index growth
In a logging or metrics pipeline, indices keep growing as new data arrives. Without automated cleanup, operators must monitor disk usage, manually delete old indices, and risk shard allocation failures when the cluster runs out of space. This reactive approach consumes operational time and can lead to performance degradation or even cluster instability.
Thesis: ILM automates rollover and deletion based on size or age
Elasticsearch Index Lifecycle Management (ILM) lets you define a policy that moves an index through phases—hot, warm, cold, delete—and triggers actions such as rollover, shrink, forceMerge, or deletion when size or age thresholds are met. By attaching the policy to an index template, every new index that matches the pattern inherits the lifecycle, removing the need for manual intervention.
Defining a simple rollover policy
The example below creates a policy that rolls over an index when it reaches 1 GB or becomes 30 days old, then moves the rolled‑over index to a delete phase after 7 days.
PUT _ilm/policy/my-logs-policy
{
"policy": {
"phases": {
"hot": {
"actions": {
"rollover": {
"max_size": "1gb",
"max_age": "30d"
}
}
},
"delete": {
"min_age": "7d",
"actions": {
"delete": {}
}
}
}
}
}
Run this request in Kibana Dev Tools or via curl with a user that has the manage_ilm cluster privilege (typically a superuser or a role with manage on index and ilm resources).
Applying the policy via an index template
Next, create an index template that uses the policy and defines an alias for write operations.
PUT _index_template/my-logs-template
{
"index_patterns": ["my-logs-*"],
"template": {
"settings": {
"number_of_shards": 1,
"number_of_replicas": 1,
"index.lifecycle.name": "my-logs-policy",
"index.lifecycle.rollover_alias": "my-logs-write"
}
}
}
Then bootstrap the first index that the alias will point to:
PUT my-logs-000001
{
"aliases": {
"my-logs-write": {
"is_write_index": true
}
}
}
All subsequent indexing requests should target the alias my-logs-write; ILM will handle creating my-logs-000002, my-logs-000003, etc., when the rollover conditions are satisfied.
Worked example: triggering a rollover
To observe the policy in action, index enough data to push the current index over the 1 GB threshold.
- Generate bulk data (adjust the number of lines to approximate 1 GB in your environment).
- Run the bulk request against the write alias:
POST my-logs-write/_bulk
{ "index": { } }
{ "message": "sample log entry", "timestamp": "2026-10-11T18:00:00Z" }
... (repeat many lines) ...
After the bulk operation, check the index list:
GET _cat/indices/my-logs-*?v&s=index
You should see a new index (e.g., my-logs-000002) with my-logs-write alias pointing to it, while the original index is no longer the write index.
Verify the lifecycle status of the old index:
GET my-logs-000001/_ilm/explain?pretty
The response should show the current phase as delete (or hot waiting to move to delete) and list any executed steps. If the cluster is low on disk, the index may stay in the hot phase; ensure sufficient free space before testing.
Trade‑off and limitation
ILM depends on the master node to execute policy steps. If the master becomes unavailable, policy progression stalls and must be resumed manually after recovery. Additionally, the rollover action requires that the cluster has enough disk space to accommodate the configured max_size; otherwise indices remain in the hot phase, causing shard allocation warnings and possible red cluster state.
A practical way to confirm the system is healthy is to monitor the cluster.health API for relocating_shards and initializing_shards and to watch the ILM explain output for errors such as INVALID_ROLLOVER_INDEX or NO_WRITE_INDEX.
Actionable closing
Start small: define a rollover policy with modest size thresholds, attach it to a template for your logging indices, and verify the first rollover with a controlled bulk load. Once the automation is trusted, adjust phases (add warm/cold with shrink or forceMerge) and retention periods to match your storage costs and query performance needs. Regularly check the ILM explain API and cluster health to catch master‑node or disk‑space issues before they affect availability.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.