Protect Backends from Slow‑Client Exhaustion with Envoy Timeouts
Slow clients can tie up backend connections and drain resources. This guide shows how to use Envoy’s HTTP Connection Manager timeouts to automatically close idle or stalled connections, protect services, and preserve latency.
28 Apr 2026, 15:14 UTC

Problem: Slow Clients Drain Backend Resources
In a microservice landscape, a single slow or malicious client can keep a TCP connection open for minutes while sending data very slowly. Because the backend service must keep the socket alive until the request completes, the process consumes memory and CPU for each idle connection. During traffic spikes, a handful of such clients can exhaust the pool of available connections, leading to request rejections, increased latency, and even service outages.
Thesis: Use Envoy’s Connection Manager Timeouts to Automate Cleanup
Envoy’s http_connection_manager exposes two key timeout knobs:
- Idle timeout – closes a connection if no data is sent or received for a specified period.
- Request timeout – terminates a request if it takes longer than a given duration.
By configuring these per‑route or per‑cluster, you can protect backend resources without touching application code, while still returning a graceful 504 Gateway Timeout to the client.
Concrete Example: A 5‑Second Idle and 30‑Second Request Timeout
Below is a minimal envoy.yaml fragment that demonstrates the configuration. Replace <YOUR_BACKEND_CLUSTER> with the name of your upstream cluster.
static_resources:
listeners:
- name: listener_0
address:
socket_address:
address: 0.0.0.0
port_value: 80
filter_chains:
- filters:
- name: envoy.http_connection_manager
typed_config:
'@type': type.googleapis.com/envoy.extensions.filters.network.http_connection_manager.v3.HttpConnectionManager
stat_prefix: ingress_http
route_config:
name: local_route
virtual_hosts:
- name: local_service
domains: ['*']
routes:
- match: { prefix: '/' }
route:
cluster: <YOUR_BACKEND_CLUSTER>
timeout: 30s # Request timeout
http_filters:
- name: envoy.router
typed_config: {}
idle_timeout: 5s # Idle timeout
# Optional: log timeout events
access_log:
- name: envoy.file_access_log
typed_config:
'@type': type.googleapis.com/envoy.extensions.access_loggers.file.v3.FileAccessLog
path: /dev/stdout
clusters:
- name: <YOUR_BACKEND_CLUSTER>
connect_timeout: 5s
type: STRICT_DNS
lb_policy: ROUND_ROBIN
load_assignment:
cluster_name: <YOUR_BACKEND_CLUSTER>
endpoints:
- lb_endpoints:
- endpoint:
address:
socket_address:
address: backend.service.local
port_value: 8080
Where to run it: Place the file in Envoy’s config directory, then restart the Envoy service. You need root or equivalent privileges to bind to port 80 and to write the log file.
Verifying the Effectiveness of the Timeouts
After deploying the configuration, perform the following checks:
- Simulate slow clients – use
aborwrkwith a very low--rateor a custom script that sends data slowly. Observe that Envoy terminates the connection after 5 s of inactivity or 30 s of request processing. - Check metrics – query the Envoy Prometheus endpoint (
http://envoy:9901/stats/prometheus) and look for spikes inhttp_connection_manager.streams_ongoingfollowed by a drop. Also inspectcluster_manager.upstream_cx_totalfor a reduction in open downstream connections. - Review logs – Envoy writes a line containing
timeoutwhen a timeout occurs. A typical entry looks like:
2026-10-10T20:40:12.123Z [info] timeout: request timeout 30s, stream 12345
- Measure backend load – before and after applying the timeouts, monitor the backend’s CPU and memory usage. A noticeable drop indicates that idle connections are being cleaned up.
Trade‑offs and Limitations
- Too short, and you break legitimate traffic – a 5‑second idle timeout may cut off a user typing a large JSON payload. Adjust per‑route if needed.
- Too long, and you lose the benefit – setting request timeouts to 5 minutes defeats the purpose; keep them conservative (e.g., 30 s for most services).
- Double‑timeout race conditions – if your backend also has a 30‑second timeout, the client might receive a
504from Envoy before the backend times out, causing confusion in downstream error handling. Coordinate timeout values. - Version differences – the exact field names (e.g.,
idle_timeout) may differ across Envoy releases. Always consult theenvoy.yamlschema for your version.
Actionable Next Steps
- Audit current connection usage: run
curl http://envoy:9901/stats | grep streams_ongoingto see baseline open connections. - Deploy the sample configuration in a staging environment and run a controlled load test with slow clients.
- Adjust
idle_timeoutandtimeoutvalues per service based on observed payload sizes and processing times. - Enable Envoy’s
access_logto capture timeout events for future debugging. - Integrate the metrics into your observability stack (Prometheus + Grafana) to alert on spikes in
streams_ongoing.
By fine‑tuning these timeout settings, you can protect your microservices from resource exhaustion caused by slow or stalled clients, while keeping user experience smooth with clear 504 Gateway Timeout responses.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.