Using Envoy’s Local Rate Limit to Shield Services from Traffic Spikes
Learn how Envoy’s built‑in local rate‑limit filter protects downstream services from traffic spikes using a token‑bucket algorithm, with a concrete configuration example and verification steps.
16 Dec 2025, 01:37 UTC

Problem: Sudden traffic bursts overwhelm a downstream API
When a client or a misbehaving service sends a burst of requests that exceeds the capacity of a backend, the backend can experience latency spikes, resource exhaustion, or even cascading failures. In many deployments adding an external rate‑limiting service introduces latency and operational complexity, especially when the goal is simply to protect a single Envoy proxy from a noisy neighbor.
How Envoy’s Local Rate Limit works
Envoy includes the envoy.filters.http.local_rate_limit filter, which implements a token‑bucket algorithm locally inside each worker thread. The filter is configured with a tokens per second refill rate and a burst capacity (the maximum number of tokens the bucket can hold). Each incoming request consumes one token; if the bucket is empty, Envoy immediately returns a 429 Too Many Requests response without contacting any external service. Because the state lives in the proxy, the decision is made with minimal latency, but the limit is per‑worker thread (or per‑instance if you run a single‑threaded Envoy). The aggregate throughput of a fleet is the sum of the limits of all instances.
Worked example: configuring a listener
The following snippet shows a minimal Envoy configuration that enables local rate limiting on an HTTP listener. Replace the placeholders with values appropriate for your environment.
static_resources:
listeners:
- name: listener_0
address:
socket_address: { address: 0.0.0.0, port_value: 10000 }
filter_chains:
- filters:
- name: envoy.filters.network.http_connection_manager
typed_config:
"@type": type.googleapis.com/envoy.extensions.filters.network.http_connection_manager.v3.HttpConnectionManager
stat_prefix: ingress_http
route_config:
name: local_route
virtual_hosts:
- name: backend
domains: ["*"]
routes:
- match: { prefix: "/" }
route: { cluster: backend_cluster }
http_filters:
- name: envoy.filters.http.local_rate_limit
typed_config:
"@type": type.googleapis.com/envoy.extensions.filters.http.local_rate_limit.v3.LocalRateLimit
stat_prefix: rate_limit
token_bucket:
max_tokens: 20 # burst size
tokens_per_fill: 10 # refill rate (tokens per second)
fill_interval: 1s
- name: envoy.filters.http.router
clusters:
- name: backend_cluster
connect_timeout: 0.25s
type: LOGICAL_DNS
lb_policy: ROUND_ROBIN
load_assignment:
cluster_name: backend_cluster
endpoints:
- lb_endpoints:
- endpoint:
address:
socket_address: { address: backend.example.com, port_value: 8080 }
To verify the behavior, you can use a load‑testing tool such as wrk from a machine with network access to the Envoy listener. The test should be run with a rate that exceeds the configured refill rate.
# Run on a test client (requires permission to open outbound connections)
wrk -t2 -c10 -d30s --rate 15 http://localhost:10000/
Envoy will initially return 200 OK while tokens are available. Once the token bucket is exhausted (after roughly 20 requests in this example), subsequent requests will receive 429 Too Many Requests. You can confirm this by inspecting the response codes in the wrk output or by checking Envoy’s access logs for the response_code_details field showing "rate_limit".
Trade‑offs and limitations
- Per‑instance scope: The limit applies to each Envoy worker thread (or each instance if you run a single‑threaded proxy). To achieve a cluster‑wide quota you must either run a single instance or sum the limits across all instances, which may lead to over‑ or under‑protection if the traffic distribution is uneven.
- Burst size risk: A large
max_tokensallows a sudden spike to pass through before the bucket is drained, potentially overwhelming a fragile backend. Choose a burst size that matches the backend’s ability to absorb short‑term surges. - No cross‑proxy coordination: Unlike the global rate‑limit filter, the local filter cannot share state with other Envoys, so it cannot enforce a strict, shared quota across a service mesh.
Actionable steps
- Determine the acceptable request rate for your downstream service (tokens per second) and the maximum burst it can tolerate.
- Add the
envoy.filters.http.local_rate_limitfilter to the listener’s HTTP filter chain, settingtokens_per_fillandmax_tokensaccordingly. - Reload Envoy (e.g.,
envoy -c /etc/envoy/envoy.yaml --restart-epoch 1) – this changes state, so keep a copy of the previous configuration for rollback if needed. - Generate test traffic that exceeds the configured rate using
wrk,locust, or similar, and verify that the response code shifts from200to429once the bucket is empty. - Monitor Envoy’s statistics (
rate_limit..local_rate_limit) and access logs to ensure the filter is behaving as expected in production.
By applying Envoy’s local rate limit you gain a lightweight, low‑latency defense against traffic spikes without introducing an external dependency, while being aware of its per‑instance nature and burst‑size considerations.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.