Search latency climbs only under concurrent clients: diagnosing the search thread pool queue
0 reputation · 20 Aug 2026, 09:58 UTC
I am trying to understand a latency pattern in Elasticsearch search requests that only appears when many clients query at once. Single-request latency looks fine, but as concurrency rises, response times grow roughly in proportion to load rather than staying flat.
My understanding is that the search thread pool uses a fixed number of threads with an effectively unbounded queue by default, so saturated threads cause requests to wait rather than be rejected. The request circuit breaker trips on heap pressure, not queue depth, so it may not protect against this kind of queueing delay. I am assuming a 7.x cluster; I know older versions allowed bounding threadpool.search.queue_size, but I am unsure what the current recommended approach is.
My goal is to confirm the queue is the bottleneck and decide between accepting queueing, bounding the queue (and handling client rejections/retries), or scaling out.
Specifically:
- Which metrics (e.g.,
_nodes/statsthread pool queue/rejected counts, p99 latency) best confirm queue-induced latency versus slow queries? - In current 7.x/8.x versions, is bounding the search queue still supported and advisable, or is there a preferred load-shedding mechanism?
- How do teams typically size the queue or shard count to balance tail latency against rejection rates?