Envoy adaptive concurrency filter: sizing limits when enforcement is per worker
0 reputation · 11 Dec 2023, 20:21 UTC
The HTTP adaptive concurrency filter (envoy.filters.http.adaptive_concurrency) with its gradient controller is a documented way to hold a service near its latency knee. The unresolved part is how to express the limit when the filter instance and its state live per Envoy worker thread.
With N workers, the effective global concurrency is roughly N times the configured or computed limit. Circuit breaker thresholds (max_connections, max_pending_requests, max_requests) and HTTP/2 max_concurrent_streams are also described as per-worker values, so a global intent divided by worker count may still be exceeded on one worker.
The gradient controller reacts to observed latency but cannot distinguish upstream queueing from its own. Field names, defaults and stat names vary by Envoy release, so any sizing decision needs confirmation against the installed version's documentation.
Questions: Should limits be configured as per-worker values or as a global intent divided by worker count? When latency samples include upstream queueing, how should the adaptive limit be trusted? Which documented stat or config field confirms per-worker enforcement in a given build?