Reducing Alert Noise for Transient Export Failures
To reduce alert noise caused by Failed to export span logs during transient network hiccups, you must shift the failure detection threshold from individual event logging to aggregated metric thresholds. Because the SDK's BatchSpanProcessor typically lacks a built-in, configurable retry policy that mirrors the Collector's retry_on_failure, the SDK will continue to log every failure it encounters before the batch is dropped.
Likely Explanation: The SDK-Collector Gap
The discrepancy occurs because the OpenTelemetry Collector is designed as a durable proxy with sophisticated buffering and retry logic, whereas the SDK BatchSpanProcessor is designed for low-overhead telemetry emission. When a network hiccup occurs, the SDK attempts to push a batch; if the request fails, the SDK logs the error and drops the batch to avoid memory leaks or blocking the application. This creates a "noisy" environment where the SDK reports a failure that the Collector might have eventually resolved had the data reached it.
Recommended Implementation Steps
Since a public API for retry intervals in the BatchSpanProcessor is not a standard feature across all language SDKs, use the following strategy to preserve visibility while silencing noise:
- Suppress Log-Based Alerting: Move away from alerting on the string
"Failed to export span" in your logging pipeline. This is a binary signal that does not distinguish between a 10ms flicker and a total outage.
- Implement Metric-Based Alerting: Monitor the SDK's internal telemetry (if available) or the Collector's
exporter_send_failed_spans metric.
- Define a Sustained Failure Threshold: Set alerts based on a rate of failure over a window (e.g.,
sum(rate(exporter_send_failed_spans[5m])) > X) rather than a single event.
- Align Collector Retries: Ensure the Collector's
retry_on_failure is configured to handle the transient spikes, allowing the SDK to remain a "fire-and-forget" mechanism.
Verification and Safety
To verify that your new alerting threshold ignores hiccups but catches outages, you can simulate a scoped failure using a network tool like tc (traffic control) on a staging environment:
# Simulate 10% packet loss to the collector endpoint
sudo tc qdisc add dev eth0 root netem loss 10%
Verify that the Failed to export span logs still appear (confirming the SDK behavior), but your operational alerts remain silent. Then, simulate a total outage (loss 100%) to ensure the aggregated metric triggers the alert.
Diagnostic Detail Needed
To provide a more specific API recommendation: Which specific language SDK (e.g., Java, Go, JS) and version are you currently utilizing? Configuration options for the BatchSpanProcessor vary significantly between implementations.