Configuring Envoy’s Rate Limit Filter for Per‑Route Traffic Control
Learn how to protect a backend service by applying Envoy’s built‑in Rate Limit Service filter to a specific route, with a concrete configuration example and operational considerations.
29 Jan 2026, 02:46 UTC

Problem: traffic spikes overwhelm your backend
When a client or a misbehaving service bursts traffic, backend instances can become saturated, leading to latency spikes or outright failures. Relying solely on application‑level throttling often means the damage is already done before the limit is enforced.
Thesis: use Envoy’s Rate Limit Service filter
Envoy provides a built‑in filter that calls an external gRPC Rate Limit Service. By evaluating descriptors such as the client IP, request headers, or dynamic metadata, the filter can decide to allow or deny a request before it reaches the upstream cluster. When correctly placed, this gives you a lightweight, configurable guardrail for any route.
How the filter works
The filter is configured as a http_filters entry that references a cluster implementing the envoy.service.ratelimit.v3.RateLimitService API. For each request, Envoy builds a set of descriptors (e.g., remote_address, header:x-api-key) and sends them to the service. If any descriptor exceeds its configured limit, the service returns an OVER_LIMIT response, Envoy returns HTTP 429, and increments the ratelimit_applied statistic.
Where to apply the filter
You can attach the filter at three levels:
- Virtual host – applies to all routes under that host.
- Route – granular control per path or virtual host + path combination.
- Weighted cluster – useful when you want different limits for different upstream subsets.
Choosing the route level lets you protect a specific API endpoint without affecting others on the same virtual host.
Worked example: per‑IP limit on a route
The following snippets show a minimal Envoy configuration that limits a route to five requests per ten seconds based on the client’s IP address. The example uses the reference implementation envoyproxy/ratelimit as the external service.
# listener.yaml
static_resources:
listeners:
- name: listener_0
address:
socket_address: { address: 0.0.0.0, port_value: 10000 }
filter_chains:
- filters:
- name: envoy.filters.network.http_connection_manager
typed_config:
"@type": type.googleapis.com/envoy.extensions.filters.network.http_connection_manager.v3.HttpConnectionManager
stat_prefix: ingress_http
route_config:
name: local_route
virtual_hosts:
- name: backend
domains: ["*"]
routes:
- match: { prefix: "/api" }
route:
cluster: backend_service
# Apply the rate limit filter only to this route
typed_per_filter_config:
envoy.filters.http.ratelimit:
"@type": type.googleapis.com/envoy.extensions.filters.http.ratelimit.v3.RateLimit
domain: envoy
descriptors:
- entries:
- key: remote_address
value: ""
rate_limit_service:
grpc_service:
envoy_grpc:
cluster_name: ratelimit_cluster
# Fail open if the service is unavailable
failure_mode_allow: true
http_filters:
- name: envoy.filters.http.router
clusters:
- name: backend_service
connect_timeout: 0.25s
type: logical_dns
lb_policy: round_robin
load_assignment:
cluster_name: backend_service
endpoints:
- lb_endpoints:
- endpoint:
address:
socket_address: { address: backend.example.com, port_value: 8080 }
- name: ratelimit_cluster
connect_timeout: 0.25s
type: logical_dns
lb_policy: round_robin
load_assignment:
cluster_name: ratelimit_cluster
endpoints:
- lb_endpoints:
- endpoint:
address:
socket_address: { address: ratelimit.example.com, port_value: 6070 }
# rate_limit_service.yaml (run separately, e.g., docker run -p 6070:6070 envoyproxy/ratelimit)
# The service must be configured with a descriptor that matches the above:
# domain: envoy
# descriptors:
# - entries:
# - key: remote_address
# value: ""
# rate_limit:
# unit: second
# requests_per_unit: 0.5 # 5 per 10 seconds
Where to run the commands:
- Start the rate‑limit service:
docker run -d -p 6070:6070 envoyproxy/ratelimit(requires ability to pull images and bind ports). - Start Envoy with the above configuration mounted as
envoy.yaml:docker run -d -p 10000:10000 -v $(pwd)/envoy.yaml:/etc/envoy/envoy.yaml envoyproxy/envoy. - Generate traffic:
for i in {1..8}; do curl -s -o /dev/null -w "%{http_code}\n" http://localhost:10000/api; done. - Expect the first five lines to be
200and the next three to be429. - Check Envoy admin stats (by default on port 9901):
curl localhost:9901/stats | grep ratelimit. You should see increments inratelimit_ok,ratelimit_applied, and possiblyratelimit_errorif the service is unreachable.
Required permissions: the user must be able to run Docker containers (or have equivalent privileges to launch Envoy and the rate‑limit service) and open the listed ports on the host.
Trade‑offs and limitations
While the filter adds protection, it introduces a round‑trip to the external gRPC service for every request that matches the descriptor set. If the service is slow or the descriptor set is overly granular (e.g., combining many high‑cardinality headers), each request may experience added latency. Mitigate this by:
- Setting a low
timeouton therate_limit_serviceblock. - Adjusting
max_callsto limit concurrent calls. - Keeping descriptors simple (e.g., only
remote_addressor a low‑cardinality header).
Availability of the rate‑limit service is critical. With failure_mode_allow: true (as in the example) Envoy will treat a service failure as “allow”, preventing a total block but possibly allowing bursts. Setting it to false will cause Envoy to deny requests when the service is unreachable, which can lead to unexpected traffic drops if the service experiences a transient outage.
Actionable next steps
- Identify the API or route that needs protection and decide on the descriptor(s) that best represent the client identity you want to limit.
- Deploy a highly available rate‑limit service (multiple replicas behind a load balancer) and verify its health endpoints.
- Apply the filter at the route level, start with a conservative limit, and monitor
ratelimit_appliedand latency via Envoy’s admin interface. - Tune the
timeoutandmax_callssettings based on observed latency and service capacity. - Plan a rollback: simply remove the
typed_per_filter_configblock from the route configuration and reload Envoy; this does not alter persistent state, so a rollback is safe.
By following these steps you can enforce per‑route traffic limits with Envoy’s built‑in Rate Limit Service, gaining a clear, observable guardrail against traffic spikes while understanding the operational trade‑offs involved.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.