Taming 5xx Cascades in Envoy with a Careful Retry Policy
Transient 5xx from an upstream can cascade into latency spikes. A bounded Envoy HTTP connection manager retry policy with specific retry_on codes, exponential backoff, and per-try timeout absorbs blips without amplifying load.
22 Feb 2026, 12:10 UTC

Upstream returns 502 for a few seconds and suddenly your service latency spikes and error rate climbs. The useful takeaway is that Envoy can absorb transient 5xx with a bounded retry policy, but only if you limit what to retry, how long to wait between attempts, and how long each attempt may run.
The failure amplification problem
When a downstream service hiccups, clients often retry immediately and in unison. Envoy sits in front of the upstream as the HTTP connection manager, the Envoy filter that terminates client HTTP and proxies to clusters. Without control, those client retries amplify load on an already unhealthy upstream, creating a thundering herd and cascading latency.
A retry policy in the route configuration lets Envoy retry a single client request internally, with backoff and a cap on attempts, before returning a final error to the client. That keeps the client from hammering you and gives a brief window for the upstream to recover.
What the retry policy actually controls
Retry configuration lives on a RouteAction and is evaluated per request. Key knobs are:
retry_on: conditions that trigger a retry, such as5xx,gateway-error,connect-failure,reset, orretriable-4xx. Envoy will not retry non-idempotent methods by default unless the request is explicitly marked retryable.num_retries: maximum number of additional attempts. Total attempts = 1 + num_retries.per_try_timeout: a timeout for each individual attempt, independent of the overall route timeout.retry_backoff: base interval and max interval for exponential backoff between attempts.
Backoff is important because immediate retries increase contention. Per-try timeout prevents a slow upstream from consuming the whole route timeout on the first try.
Route-specific example with bounded retries
The following static configuration shows a route that retries only on specific 5xx gateway errors with exponential backoff and a per-try timeout. Replace placeholders with your own cluster and host values.
static_resources:
listeners:
- name: http_listener
address:
socket_address: { address: 0.0.0.0, port_value: 8080 }
filter_chains:
- filters:
- name: envoy.filters.network.http_connection_manager
typed_config:
'@type': type.googleapis.com/envoy.extensions.filters.network.http_connection_manager.v3.HttpConnectionManager
stat_prefix: ingress_http
route_config:
name: local_route
virtual_hosts:
- name: backend
domains: ['*']
routes:
- match: { prefix: '/' }
route:
cluster: backend_cluster
retry_policy:
retry_on: \\\"502,503,504\\\"
num_retries: 2
per_try_timeout: 0.5s
retry_backoff:
base_interval: 0.025s
max_interval: 0.25s
http_filters:
- name: envoy.filters.http.router
This means: retry up to 2 additional times, only on 502, 503, or 504, wait ~25ms then ~50ms before the next attempt, capped at 250ms, and abort any single attempt after 500ms. The settings are route-specific, so you can be more conservative for read-only paths and stricter for writes.
Run Envoy with the config and point traffic at a test upstream. You can observe retry behavior via the admin interface at http://localhost:9901/stats. Look for counters under http.router and cluster.backend_cluster that track retry attempts and upstream responses. A practical check is to make the upstream return 502 intermittently and confirm the client sees a single request while Envoy performs the configured number of internal attempts.
Trade-offs and safety limits
Aggressive retries improve success rates for transient errors but increase load on unhealthy upstreams and can mask real problems. Combine retries with outlier detection and circuit breakers so Envoy stops sending traffic to consistently failing hosts.
Idempotency matters. Retrying POST, PUT, or PATCH can cause duplicate side effects unless the application guarantees idempotency or you explicitly allow retries via headers. Envoy by default avoids retrying non-idempotent methods for 5xx unless the request is marked retryable.
Do not rely on retries alone. Use them as a short-term shock absorber, keep num_retries low, use backoff, and set per_try_timeout shorter than the overall route timeout. Monitor retry rates and upstream latency to detect when retries are hurting more than helping.
Close the loop by defining a clear retry budget per route, validating it with controlled fault injection, and pairing it with circuit breaking. That gives resilience without turning a blip into a cascade.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.