Choosing an Envoy Load Balancing Policy: Decision Guide for Service-to-Service Traffic
Compare Envoy's built-in load balancing algorithms (Round Robin, Least Request, Ring Hash, Maglev, Random) for service-to-service traffic. Learn trade-offs, session affinity, failover behavior, and how to validate your choice in Envoy 1.28+.
16 Apr 2026, 00:43 UTC

The Decision: Pick One Policy Per Cluster
Envoy selects a load balancing policy per upstream cluster via the lb_policy field. The choice determines how requests are distributed across healthy endpoints, whether session affinity is provided, and how the cluster reacts to endpoint changes. Constraints include:
- Need for request affinity (stateful gRPC streams, WebSocket)
- Heterogeneous endpoint capacity or request latency
- Cluster size (tens to thousands of endpoints)
- Frequency of endpoint membership changes (e.g., Kubernetes pod churn)
- Zone-aware routing requirements
All policies respect health checks, outlier detection, and zone-aware routing. Zone affinity can be enabled per cluster to prefer local-zone endpoints before crossing AZ boundaries.
Quick Comparison
| Algorithm | Best For | Session Affinity | Load Adaptation | Scale Suitability | Config Complexity |
|---|---|---|---|---|---|
| Round Robin (default) | Simple, homogeneous workloads | No | None (ignores load variance) | Any | Low |
| Least Request | General-purpose RPC/HTTP, heterogeneous workloads | No | Picks two random endpoints, sends to the one with fewer active requests (P2C) | Any | Low (set choice_count: 2) |
| Ring Hash | Stateful protocols requiring affinity | Yes (consistent hashing) | No | Up to ~1000 endpoints | Medium (requires hash_policy) |
| Maglev | Large clusters (>1000 endpoints) needing affinity | Yes (consistent hashing, faster lookup) | No | Thousands of endpoints | Medium (requires hash_policy, CPU-intensive table generation) |
| Random | Avoiding synchronization spikes at startup | No | Statistically equivalent to Round Robin over time | Any | Low |
Trade-offs and Cautions
Least Request (Recommended Default)
Least Request with choice_count: 2 adapts to heterogeneous workloads and reduces tail latency without per-endpoint state. In Envoy 1.26+, active_request_bias can weight by latency for highly skewed request costs. It works well for most service-to-service RPC and HTTP traffic.
Ring Hash vs. Maglev for Affinity
Both provide session affinity via consistent hashing. Ring Hash uses a configurable ring size (default 1024 virtual nodes per host). Maglev uses a permutation table for O(1) lookup and better distribution at scale. Choose Maglev for clusters exceeding ~1000 endpoints. Both require a hash_policy (header, cookie, or connection properties) to define the affinity key. Hash policies using headers or cookies expose affinity to client manipulation; prefer connection properties (source IP, TLS peer cert) for untrusted clients.
Endpoint Churn and Reshuffling
Ring Hash and Maglev reshuffle on endpoint membership changes. Control virtual node count with min_ring_size and max_ring_size to limit disruption. Maglev table generation is CPU-intensive on cluster updates; avoid frequent endpoint churn (e.g., per-pod IPs in large Kubernetes clusters) or increase table_size incrementally.
Zone-Aware Routing
Zone affinity prefers local-zone endpoints. Perfect balance requires identical endpoint counts per zone; otherwise cross-zone traffic may exceed cross_zone_fallback_ratio.
Deprecated Policy
Envoy 1.28 deprecated lb_policy: CLUSTER_PROVIDED in favor of explicit policy names. Update configs before upgrading.
Concrete Implementation Examples
Least Request Cluster (General Purpose)
clusters:
- name: backend_service
type: EDS
lb_policy: LEAST_REQUEST
least_request_lb_config:
choice_count: 2
active_request_bias: 1.0 # optional, Envoy 1.26+
health_checks:
- timeout: 5s
interval: 10s
unhealthy_threshold: 3
healthy_threshold: 2
http_health_check:
path: /healthz
outlier_detection:
consecutive_5xx: 5
interval: 30s
base_ejection_time: 30s
# Zone-aware routing (optional)
locality_lb_policies:
- zone: us-east-1a
lb_policy: LEAST_REQUEST
- zone: us-east-1b
lb_policy: LEAST_REQUEST
Ring Hash Cluster (Affinity Required)
clusters:
- name: stateful_service
type: EDS
lb_policy: RING_HASH
ring_hash_lb_config:
min_ring_size: 1024
max_ring_size: 8192
hash_policy:
- header:
header_name: x-session-id
# For untrusted clients, use connection properties instead:
# connection_properties:
# source_ip: true
health_checks:
- timeout: 5s
interval: 10s
unhealthy_threshold: 3
healthy_threshold: 2
http_health_check:
path: /healthz
Replace x-session-id with the header or cookie that carries your session identifier. For Maglev, change lb_policy: MAGLEV and add maglev_lb_config: with table_size (default 65537, must be prime).
Validation Steps
1. Confirm Active Policy via Admin Interface
Run on the Envoy host (requires access to admin port, typically 9901):
curl -s http://localhost:9901/clusters?format=json | jq '.[] | select(.name=="backend_service") | {name, lb_policy, load_balancing_policy}'
Expected output shows lb_policy" : "LEAST_REQUEST" (or the policy you set). Risk: exposing admin port publicly; restrict to localhost or trusted network.
2. Synthetic Load Test
From a test client with hey or wrk installed:
hey -n 10000 -c 50 http:///service/endpoint
Then check request distribution:
curl -s http://localhost:9901/stats?filter=cluster.backend_service.upstream_rq* | grep upstream_rq
Least Request should show roughly even counts across endpoints; Round Robin will show sequential pattern. Risk: load test may affect production traffic; run against a staging cluster.
3. Trace Affinity Adherence (Ring Hash/Maglev)
Enable access logging with upstream host and request ID:
access_log:
- name: envoy.access_loggers.file
typed_config:
"@type": type.googleapis.com/envoy.extensions.access_loggers.file.v3.FileAccessLog
path: /var/log/envoy/access.log
log_format:
text_format: "%START_TIME% %REQ(:METHOD)% %REQ(X-REQUEST-ID)% %UPSTREAM_HOST% %RESPONSE_CODE%\n"
Inspect logs to verify that requests with the same hash key consistently route to the same upstream host.
Limitations and Practical Checks
- Zone imbalance: If zones have unequal endpoint counts, cross-zone traffic may exceed the fallback ratio. Monitor
cluster..upstream_cx_cross_zonestats. - Maglev CPU spikes: On large clusters with frequent endpoint updates, watch CPU usage during cluster reloads. Consider increasing
table_sizegradually. - Hash policy security: Header/cookie based affinity can be spoofed. Use
connection_propertieswithsource_iportls_peer_certfor untrusted clients. - Deprecated config: Validate config with
envoy -c config.yaml --mode validatebefore deployment to catchCLUSTER_PROVIDEDusage.
Practical verification: after deployment, monitor /stats for upstream_rq_pending_active, upstream_rq_per_try_timeout, and upstream_cx_destroy_remote_active_rq to ensure the policy behaves as expected under real traffic.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.