Diagnosing and Fixing Envoy Outlier Detection Storms
When Envoy starts ejecting hosts, 503s surge. This guide walks you through diagnosing the cause, tuning outlier detection, verifying fixes, and knowing when to involve SREs.
01 Oct 2026, 09:36 UTC

Recognize the Symptom
When a cluster starts ejecting hosts, clients see a sudden spike in 503 Service Unavailable responses. In Envoy logs you’ll see entries such as:
2026-10-10T09:00:00Z info envoy: Ejected host 10.0.0.5 from cluster myservice due to consecutive_5xx
These ejections are intentional: Envoy removes a host from load‑balancing when it thinks the host is unhealthy. If too many hosts are ejected, the cluster may become effectively empty, causing the 503 surge.
Common Root Causes
Outlier detection can trigger for several reasons. The most frequent triggers are:
- Consecutive 5xx errors – a host returns many errors in a row.
- Low success rate – the ratio of successful requests falls below a threshold.
- High request volume with failures – a sudden traffic burst reveals a flaky host.
- Detection interval mis‑set – the interval is too short, causing quick ejections.
- Upstream flapping – the host repeatedly goes up and down, confusing Envoy.
Diagnostic Checklist
Follow this ordered checklist to pinpoint the cause.
- Confirm the spike
- Check
/statsforoutlier_ejected_totalandcluster_name.upstream_rq_503.curl -s http://localhost:9901/stats?format=json | jq '.values | map(select(.name | contains("outlier_ejected_total")))' - Verify the 503 rate has increased beyond baseline (e.g., >5% of traffic).
- Check
- Identify ejection reasons
- Search logs for "Ejected host" and note the reason field.
grep "Ejected host" /var/log/envoy.log | tail -n 20 - Common reasons:
consecutive_5xx,success_rate_threshold,request_volume,detection_interval.
- Search logs for "Ejected host" and note the reason field.
- Inspect cluster configuration
- Dump the current config and locate the
outlier_detectionblock.curl -s http://localhost:9901/config_dump | jq '.dynamic_config.dump.clusters[] | select(.name=="myservice") | .outlier_detection' - Take note of
consecutive_5xx,success_rate_threshold,request_volume,detection_interval, andmin_request_volume.
- Dump the current config and locate the
- Correlate traffic patterns
- Compare Envoy metrics (e.g.,
cluster_name.upstream_rq_total) with upstream service logs to see if a traffic burst coincides with the ejection.
- Compare Envoy metrics (e.g.,
Tuning Strategies
Adjust the outlier detection settings only after you understand the root cause. Below are the most common adjustments and when to apply them.
| Root Cause | Diagnostic Indicator | Recommended Fix |
|---|---|---|
| Consecutive 5xx errors | Logs show consecutive_5xx ejection reason. | Increase consecutive_5xx threshold or patch the upstream service to reduce errors. |
| Low success rate | Logs show success_rate_threshold reason; cluster_name.upstream_rq_success < 90%. | Raise success_rate_threshold (e.g., from 90% to 95%) or increase min_request_volume to avoid small‑sample swings. |
| High request volume with failures | Logs show request_volume reason; cluster_name.upstream_rq_total spikes. | Increase request_volume threshold or extend detection_interval to give the host time to recover. |
| Detection interval too short | Frequent ejections within seconds. | Set detection_interval to at least 10s or more, depending on traffic patterns. |
| Upstream flapping | Host status toggles between healthy/unhealthy in logs. | Enable interval_jitter_percent or add a failure_mode to ignore transient failures. |
Example: Raising the 5xx Threshold
Assume the cluster myservice is ejecting on consecutive_5xx. The current config is:
outlier_detection:
consecutive_5xx: 5
interval: 10s
base_ejection_time: 30s
max_ejection_percent: 50
To reduce false positives, change consecutive_5xx to 10:
outlier_detection:
consecutive_5xx: 10
interval: 10s
base_ejection_time: 30s
max_ejection_percent: 50
Apply the new config via a hot reload or restart. After deployment, verify that outlier_ejected_total stops rising and 503 rates return to baseline.
Verification
After making changes, perform the following checks:
- Stats trend
- Run
/statsover a 5‑minute window and confirmoutlier_ejected_totalis flat.
- Run
- Error rate
- Verify
cluster_name.upstream_rq_503is back to normal (<10% of total).
- Verify
- Host availability
- Check
/healthcheckor access logs to ensure previously ejected hosts are receiving traffic again.
- Check
- Latency
- Monitor request latency; a sudden drop indicates the cluster is healthy.
Escalation Criteria
If after tuning the ejection rate remains high or the cluster is still unable to serve traffic, involve the SRE team. Escalate when:
- More than 5% of requests return 503 for over 5 minutes.
- Ejections continue despite threshold adjustments.
- Upstream service health metrics (e.g., CPU, memory, error logs) indicate degradation.
- Potential cascading failures are observed in downstream services.
During escalation, share:
- Current outlier detection config.
- Recent
outlier_ejected_totalandupstream_rq_503metrics. - Upstream service logs showing error spikes.
- Any recent traffic bursts or deployment events.
Best Practices
- Use
runtimeoverrides only for temporary experiments; persist changes in the static config. - Set
min_request_volumehigh enough to avoid reacting to a handful of requests. - Combine outlier detection with
circuit_breakersandload_sheddingfor layered resilience. - Monitor
outlier_ejected_totalcontinuously; set alerts for sudden increases. - Document any threshold changes and the rationale to aid future debugging.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.