Using Envoy Outlier Detection to Protect Canary Releases
Learn how Envoy’s outlier detection automatically removes flaky upstream hosts during canary rollouts, with a practical configuration example, verification steps, and trade‑off guidance.
19 Jul 2026, 13:00 UTC

Problem: flaky upstream can destabilize a canary rollout
When you shift a small fraction of traffic to a new version, any upstream instance that starts returning 5xx errors can cause retries, timeouts, and increased latency for those users. If the bad host stays in the load‑balancing pool, the failure rate can climb and undermine the safety guarantees of the canary.
Thesis: Envoy’s outlier detection automatically removes misbehaving hosts
Envoy can monitor each upstream host for consecutive 5xx responses or a rising error rate and temporarily eject it from the load‑balancing pool. The ejected host is allowed back after a configurable cool‑down period, giving the service time to recover without manual intervention.
How outlier detection works
Outlier detection runs per cluster. For each host Envoy tracks:
consecutive_5xx– number of successive 5xx responses before ejection.success_rate– ratio of successful requests over a sliding window; ejection triggers when the ratio falls below a threshold.base_ejection_time– initial duration a host is ejected; it can be lengthened with repeated ejections.max_ejection_percent– upper bound on how many hosts in the cluster may be ejected simultaneously.
When a host meets either the consecutive‑5xx or success‑rate condition, Envoy increments an ejection counter, stops sending traffic to that host, and starts a timer. After the timer expires, the host is reintroduced.
Configuration example
The following Envoy snippet shows a cluster named frontend with outlier detection tuned for a canary that expects at most 2% error rate.
static_resources:
clusters:
- name: frontend
connect_timeout: 0.25s
type: LOGICAL_DNS
lb_policy: ROUND_ROBIN
load_assignment:
cluster_name: frontend
endpoints:
- lb_endpoints:
- endpoint:
address:
socket_address: { address: service-a.example.com, port_value: 8080 }
outlier_detection:
consecutive_5xx: 3
interval: 10s
base_ejection_time: 30s
max_ejection_percent: 10
enforcing_consecutive_5xx: true
enforcing_success_rate: true
success_rate_minimum_hosts: 5
success_rate_request_volume: 100
success_rate_stdev_factor: 1900 # ~1.96 sigma
Place this block inside your Envoy bootstrap or static resources file. Adjust consecutive_5xx, base_ejection_time, and max_ejection_percent to match your service’s SLA and traffic volume.
Worked example: canary safety with a flaky upstream
Imagine a canary deployment where 5% of traffic goes to service-a:v2. One instance of service-a begins returning 5xx after every third request due to a bug.
- Envoy receives requests for the canary.
- The faulty instance hits the
consecutive_5xx: 3threshold. - Envoy ejects that host for
base_ejection_time: 30s. - During the ejection, the remaining healthy instances serve the canary traffic; retries (if configured) attempt other hosts, keeping latency low.
- After 30 s the host is allowed back; if it continues to fail, the ejection time may be doubled on subsequent events.
This automatic removal prevents the bad instance from dragging down the overall error rate while giving the team time to roll back or fix the bug.
Trade‑offs and limitations
Over‑aggressive settings can eject too many hosts, especially during bursty traffic spikes, leading to traffic starvation or increased latency if the remaining pool cannot handle the load. Conversely, overly permissive thresholds may let failing hosts stay too long, degrading user experience.
Outlier detection assumes relatively steady request patterns. Rapid traffic surges can cause false positives because the success‑rate window may not have enough samples. Administrators should watch ejection events and tune the interval and volume parameters accordingly.
Verifying the behavior
Enable the Envoy admin interface (e.g., admin: { address: { socket_address: { address: 0.0.0.0, port_value: 9901 } } }) and query the stats endpoint:
curl http://localhost:9901/stats | grep outlier_detection
Look for counters such as:
cluster.frontend.outlier_detection.ejection_consecutive_5xxcluster.frontend.outlier_detection.ejection_success_ratecluster.frontend.outlier_detection.ejection_active
When the faulty instance starts returning 5xx, you should see the ejection counters increase and the active ejection gauge rise to 1 (or more, depending on max_ejection_percent). After the ejection period, the active gauge should drop back to 0.
For a local test, you can run a Docker Compose stack that includes Envoy and a simple mock upstream (e.g., a Go or Python HTTP server) that returns 5xx after a configurable number of requests. Compose the services, generate traffic with a tool like hey or wrk, and observe the admin stats as described.
Actionable closing
To use outlier detection effectively:
- Start with conservative values (e.g.,
consecutive_5xx: 5,base_ejection_time: 15s,max_ejection_percent: 10). - Monitor ejection counters via the admin
/statsendpoint. - Adjust thresholds based on observed error rates and traffic patterns, ensuring the ejected host percentage stays well below
max_ejection_percentduring normal operation. - Combine outlier detection with retries and circuit breakers for layered resilience.
By letting Envoy automatically handle misbehaving hosts, you gain a safety net for canary releases and reduce the need for manual intervention when upstream instances falter.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.