Envoy circuit breaker: intermittent upstream connection pool exhaustion under aggregate limits
0 reputation · 20 Oct 2023, 03:49 UTC
0 reputation · 20 Oct 2023, 03:49 UTC
Envoy bounds upstream connection pressure through per-cluster, per-priority circuit-breaker limits such as max_connections, max_pending_requests, max_requests, and max_retries. When a limit is reached, Envoy rejects or queues work instead of opening more connections. Intermittent exhaustion appears in cluster-level counters: upstream_cx_overflow (connection limit hit), upstream_rq_pending_overflow (pending-request limit hit), and upstream_rq_retry_overflow, alongside gauges like upstream_cx_active and upstream_rq_pending_active.
max_requests_per_connection and the HTTP/2 max_concurrent_streams setting, forcing connection churn that may masquerade as a connection-count problem.Which binding constraint—connection count (max_connections) or concurrent stream/request capacity (max_pending_requests, max_requests_per_connection, HTTP/2 max concurrent streams)—is the primary driver of the observed overflow pattern? Does the per-priority exhaustion behavior warrant explicit per-priority circuit-breaker overrides, or should the aggregate limits be restructured? What controlled load-test methodology best isolates aggregate-limit exhaustion from per-host or upstream-imposed limits?
A thoughtful contribution can make all the difference. Be the first to share one.
Use comments to ask for clarification. Post a solution as an answer.
No question comments on this page.