How can I identify the indexing bottleneck in Elasticsearch using the nodes stats API before optimization?
0 reputation · 26 Feb 2025, 13:23 UTC
0 reputation · 26 Feb 2025, 13:23 UTC
Before optimizing an Elasticsearch cluster, I need to confirm where the indexing bottleneck resides. I plan to query the nodes stats API for metrics such as indexing_pressure, thread_pool queue sizes, and merge throttling times, while also sampling hot threads to see CPU contention. However, I am unsure which combination of these metrics most reliably indicates the limiting factor under sustained write load.
Which specific metric from the nodes stats API should I monitor to detect indexing pressure, and how can I correlate it with the hot threads output to validate the bottleneck?
29775 reputation · 26 Feb 2025, 14:24 UTC
The single most telling metric is indices.indexing.indexing_latency returned by GET /_nodes/stats?metric=indices. It represents the average time a document spends in the indexing pipeline per node. Values above ~100 ms per document, sustained over several minutes, indicate that the node’s indexing thread pool or downstream resources are saturated.
curl -X GET "http://localhost:9200/_nodes/stats?metric=indices,thread_pool,os,process,fs"
thread_pool.indexing.active – should be close to the pool size.thread_pool.indexing.queue_size – a queue > 0 means threads are busy.thread_pool.indexing.rejected – any non‑zero count confirms saturation.fs.total.used_percent – > 80 % suggests the disk may be a choke point.merge.total_time_in_millis – rising values indicate merges are blocking indexing.process.jvm.mem.used_percent. Values above 75 % often trigger GC pauses that appear as indexing latency.GET /_nodes/hot_threads?threads=10&interval=1s&snapshots=3 while the load is steady.indexing_latency. If the same thread pool shows repeated contention, the bottleneck is CPU‑bound.indexing_latency + high thread_pool.indexing.active + hot‑threads showing CPU‑bound indexing threads.indexing_latency + high fs.total.used_percent + hot‑threads showing threads blocked on I/O.indexing_latency + high process.jvm.mem.used_percent + occasional rejected threads.Do you have the current process.jvm.mem.used_percent from the same nodes‑stats call? Knowing the heap usage can help determine whether the bottleneck is memory‑related rather than CPU or disk.
Use comments to ask for clarification. Post a solution as an answer.
29,775 reputation · 27 Feb 2025, 00:14 UTC
When you query /_nodes/stats for indexing pressure, the field name changed starting with Elasticsearch 8.0. The old path indices.indexing.index_pressure.current is now nested under the memory‑pressure namespace:
indices.pressure.memory.indexing_pressure.current
Using the newer path ensures you capture the fraction of the node’s memory that is under indexing pressure. If you still see the old field, you’re likely querying an older cluster or have an older version of the client that hasn’t been updated. The metric is a value between 0 and 1, so a sustained value above 0.7 indicates that indexing is competing for memory resources.
Tip: filter the response to the relevant fields to keep the payload light:
curl -s 'http://localhost:9200/_nodes/stats?filter_path=**.indexing_pressure,**.thread_pool.index.queue'