Query concurrency limit change in InfluxDB 2.7 leads to higher latency for complex Flux queries
0 reputation · 10 Jun 2021, 06:14 UTC
0 reputation · 10 Jun 2021, 06:14 UTC
When the query-concurrency-limit setting is enabled in InfluxDB 2.7 to restrict simultaneous Flux query execution, users report increased latency for certain analytical queries that previously completed within sub‑second times.
Determine whether the observed latency increase stems from the query scheduler hitting the concurrency limit or from heightened storage I/O caused by the limited parallelism.
It is unclear which internal metrics best reflect scheduler contention versus disk pressure, and how to isolate the effect without altering other configuration parameters.
Which specific counters in the _internal monitoring system indicate that the query scheduler is the bottleneck?
How can storage‑related metrics be distinguished from scheduler‑related metrics when concurrency is limited?
What reproducible workload characteristics (e.g., query shape, time range, cardinality) should be used to isolate the impact of the query-concurrency-limit setting?
26525 reputation · 10 Jun 2021, 14:33 UTC
query-concurrency-limit is reached, new queries are queued before execution begins. To identify if the scheduler is the bottleneck, monitor the following internal metrics (typically via the _internal database or the /metrics endpoint):
influxdb_fluxquery_queued_duration: A high value here indicates queries are waiting for an available slot in the concurrency pool.influxdb_query_active_count: If this consistently sits exactly at your configured concurrency limit, the scheduler is actively throttling requests.flux_task_execution_duration: While specific to tasks, high spikes here alongside queued queries suggest the engine is saturated.influxdb_storage_engine_read_latency: Increased latency in this metric indicates that the disk subsystem cannot keep up with the read requests from the Flux engine.influxdb_storage_read_operations_total: If this peaks while latency rises, the storage engine is struggling with the volume of data retrieval.query-concurrency-limit, use the following reproducible workload characteristics:
filter() on a unique tag) against a "heavy aggregation" (e.g., aggregateWindow() over a large time range). If only the aggregation slows down under load, it is CPU/Memory contention.Note: Use the influx.profile() function within your Flux scripts to see which specific operators (e.g., joins, pivots) are consuming the most clock time during execution.
Use comments to ask for clarification. Post a solution as an answer.
26,525 reputation · 10 Jun 2021, 09:20 UTC
The influxdb_fluxquery_queued_duration metric reports how long a query waits before a slot becomes available. It is expressed in seconds and is aggregated as a histogram. A persistent spike (e.g., >0.5 s) while influxdb_query_active_count sits at the configured limit indicates that the scheduler is the bottleneck, not storage I/O.
To isolate the effect, compare the queued duration against influxdb_storage_engine_read_latency. If the read latency remains low but queued duration is high, you’re seeing scheduler contention. If both rise together, the disk subsystem is likely the culprit.
Remember that the default OSS limit is 0 (unlimited). Enterprise or patched releases may set a positive cap, so always verify with influxd config get query.max-concurrent-queries after any upgrade.