Diagnosing Envoy Outlier Detection Ejections: A Step‑by‑Step Guide
A concise diagnostic guide for Envoy outlier detection: spot ejection symptoms, check counters and config, adjust thresholds, and know when to escalate.
14 May 2026, 02:13 UTC

Recognizable Condition
You notice increased latency or error rates for a service, and Envoy logs show messages like "[outlier detection] host ejected" or the admin endpoint reports rising ejection_* counters for a cluster. Traffic to one or more upstream hosts drops suddenly while other hosts appear healthy.
Cause/Diagnostic Table
| Symptom | Likely Cause | Key Indicator |
|---|---|---|
| Sudden drop in requests to a specific host | Host marked as outlier and ejected | Increase in cluster..outlier_detection.ejection_eject counter |
| Elevated latency across the pool | Success‑rate or latency threshold too low, causing premature ejection | High success_rate or success_rate_requests_total with low success ratio |
| No ejection events but errors persist | Outlier detection disabled or mis‑configured | Zero ejection_* counters; outlier_detection block missing or disabled |
Ordered Checks
- Verify admin access
Ensure you can reach Envoy’s admin interface (default port
9901) from your workstation or a bastion host. You need read‑only permission on the/statsendpoint.curl -s http://localhost:9901/stats | grep outlier_detectionLook for counters such as
cluster.my_service.outlier_detection.ejection_ejectandcluster.my_service.outlier_detection.success_rate. - Check recent ejection events
If the counter is non‑zero, note the timestamp (Envoy exports a histogram
ejection_timethat can be queried viastatswith the:valuesuffix).curl -s http://localhost:9901/stats | grep ejection_eject - Inspect the cluster’s outlier detection configuration
Retrieve the current configuration via xDS or the static file. For a static config, examine the
outlier_detectionblock:cluster: name: my_service outlier_detection: interval: 10s base_ejection_time: 30s max_ejection_percent: 10 success_rate_minimum_hosts: 5 success_rate_request_volume: 100 success_rate_stdev_factor: 1900 # ~19.0% success rate threshold consecutive_gateway_failure: 5 consecutive_5xx: 5If using xDS, run:
curl -s http://localhost:9901/config_dump | jq '.configs[].dynamic_clusters[] | select(.name=="my_service") | .outlier_detection' - Correlate with application metrics
From the upstream service, verify error rates and latency during the same window. If the upstream is healthy but Envoy still ejects, the thresholds are likely too aggressive.
- Test with controlled failure
In a staging environment, use a tool like
heyorwrkto generate a known number of 5xx responses or added latency to a single host, then watch the ejection counter increase.hey -z 30s -m GET -h "Host: upstream-host" http://envoy-admin:9901/...Observe whether ejection occurs after the configured
consecutive_gateway_failureorsuccess_ratethreshold.
Fixes Tied to Findings
- Host truly unhealthy
If the upstream is returning errors or latency spikes, fix the root cause (application bug, resource exhaustion). Once the host passes health checks, Envoy will automatically stop ejecting it after the
base_ejection_timeelapses. - Thresholds too low
Increase
success_rate_stdev_factor(or the equivalentsuccess_ratevalue) to require a larger deviation before ejection. For TCP, raiseconsecutive_connection_failureor increasebase_ejection_timeto reduce churn.Example adjustment:
outlier_detection: success_rate_stdev_factor: 2100 # ~21% success rate threshold base_ejection_time: 60s - Feature disabled
Ensure the
outlier_detectionblock is present and not set to{}(which disables the feature). Add or restore the block with sensible defaults.
Escalation Criteria
- Ejection counters continue to rise after threshold adjustments and host health is confirmed.
- Multiple clusters exhibit simultaneous ejection storms, suggesting a mis‑configured global runtime or xDS delivery issue.
- Admin endpoint shows
outlier_detectioncounters stuck at zero while upstream logs indicate failures—possible configuration not being loaded (checkconfig_dumpfor stale version). - In these cases, collect Envoy logs (
--log-level debug) and the full configuration dump, then engage the platform team to verify xDS version consistency and restart Envoy if needed.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.