OpenTelemetry Collector Tail Sampling: Reduce Ingestion Costs While Keeping Errors and Slow Traces
Use the OpenTelemetry Collector's Tail Sampling Processor to keep errors and slow traces while probabilistically sampling the rest. This guide covers a working configuration, policy ordering, memory limits, and common placement mistakes.
24 May 2026, 00:18 UTC

The Problem: Full-Fidelity Tracing Is Expensive
Sending every trace to your observability backend works fine at low volume. At scale, ingestion costs and storage retention become the primary constraints. Head-based sampling (deciding at trace start) is cheap but blind—you cannot preferentially keep errors or high-latency requests because you don't know the outcome yet.
Use the Tail Sampling Processor when you need to sample after a trace completes, so you can apply policies like "always keep errors," "keep traces slower than 500 ms," and "probabilistically sample the rest." This runs in the OpenTelemetry Collector, not in your application code.
How Tail Sampling Works
The tail_sampling processor buffers spans by trace ID until one of two conditions occurs:
- The trace appears complete (all spans have ended), detected via span end signals.
- A configurable
wait_timeelapses (default 30 s), forcing a decision on potentially incomplete traces.
Once the buffer window closes, the processor evaluates policies in order. The first matching policy decides: sample or drop. If no policy matches, the trace is dropped. Policies include:
latency— sample traces exceeding a duration threshold.status_code— sample traces with a specific status (e.g., ERROR).probabilistic— sample a percentage of remaining traces.string_attribute,numeric_attribute— sample based on span attributes.
The processor stores trace data in memory by default. For multi-instance collectors, an optional Redis backend shares state, but adds network latency and operational complexity.
Worked Configuration Example
Below is a minimal, production-style collector configuration. Save it as collector-config.yaml and run the collector with --config=collector-config.yaml. Requires a collector binary v0.100.0 or later.
receivers:
otlp:
protocols:
grpc:
http:
processors:
# Bound memory before tail sampling sees the data
memory_limiter:
check_interval: 1s
limit_mib: 512
spike_limit_mib: 128
# Batch spans to reduce downstream calls; MUST precede tail_sampling
batch:
timeout: 10s
send_batch_max_size: 1024
tail_sampling:
# Total traces held in memory at once
num_traces: 50000
# Wait longer than your p99 trace duration
wait_time: 30s
policies:
- name: errors-always
type: status_code
status_code:
status_codes: [ERROR]
- name: slow-requests
type: latency
latency:
threshold_ms: 500
- name: baseline-10pct
type: probabilistic
probabilistic:
sampling_percentage: 10
exporters:
otlp:
endpoint: tempo:4317
tls:
insecure: true
service:
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter, batch, tail_sampling]
exporters: [otlp]
Key Placement Rule
Place tail_sampling after batch (and any other span-aggregating processors) but before exporters. If you put it before batch, each span is evaluated individually, breaking trace assembly and defeating the purpose.
Policy Order Matters
The example orders policies from most specific to least specific:
- errors-always — catches any trace with an ERROR status code, regardless of latency.
- slow-requests — catches traces > 500 ms that didn't already match errors.
- baseline-10pct — probabilistically samples 10 % of everything else.
Reversing this order would cause the probabilistic policy to consume traces before the error/latency policies ever see them.
Operational Limits and Trade-offs
Memory and CPU
The processor holds complete trace data in memory until wait_time expires. With num_traces: 50000 and average trace size of 2 KB, expect ~100 MB baseline plus overhead. Monitor the collector's process_memory_bytes and the processor's traces_dropped metric (exposed via Prometheus). If traces_dropped rises, increase num_traces or scale horizontally.
Added Latency
Traces that complete before wait_time are released immediately. Incomplete traces wait the full wait_time. Set wait_time to your 99th-percentile trace duration plus a safety margin (e.g., if p99 is 12 s, use 30 s). Longer wait times increase memory pressure; shorter times risk sampling incomplete traces.
Multi-Instance State
Without Redis, each collector instance makes independent decisions. The same trace ID arriving at different instances (due to load balancing) will be evaluated separately, potentially causing inconsistent sampling. Redis backend (policy_evaluator: redis in config) solves this but introduces a network hop and a single point of failure—run Redis with HA (Sentinel or Cluster) if you go this route.
Common Mistakes
| Mistake | Symptom | Fix |
|---|---|---|
tail_sampling before batch | Traces never sampled; traces_received increments but traces_sampled stays zero | Move tail_sampling after batch in the pipeline |
wait_time too short (e.g., 5 s) | Long-running traces sampled incompletely; missing root spans | Set wait_time > p99 trace duration |
| Probabilistic policy first | Errors and slow traces dropped at the probabilistic rate | Order policies: specific → general |
No memory_limiter | Collector OOM-killed under load | Add memory_limiter before tail_sampling |
| Redis backend without HA | Collector stalls when Redis node fails | Deploy Redis Sentinel/Cluster; tune redis_timeout |
Verification Steps
- Deploy the collector with the example config to a test environment.
- Generate synthetic traces:
# Run from a host with network access to the collector otel-cli span --service test-svc --name "fast-ok" --duration 50ms otel-cli span --service test-svc --name "slow-ok" --duration 800ms otel-cli span --service test-svc --name "error-case" --status error --duration 100ms - Query your backend (Tempo, Jaeger, etc.) and confirm:
fast-okappears ~10 % of the time.slow-okappears 100 % of the time.error-caseappears 100 % of the time.
- Check collector metrics (Prometheus endpoint, usually
:8888/metrics):
Verify the counters align with your policy expectations.otelcol_processor_tail_sampling_traces_received_total otelcol_processor_tail_sampling_traces_sampled_total otelcol_processor_tail_sampling_traces_dropped_total
When Not to Use Tail Sampling
- Throughput exceeds what a single collector instance can buffer (consider head sampling in the SDK instead).
- Your traces are extremely long-running (minutes), making
wait_time impractically large. - You cannot tolerate the added tail latency for incomplete traces.
Summary Checklist
- Place
memory_limiter→batch→tail_samplingin the traces pipeline. - Set
wait_time> p99 trace duration. - Order policies: status_code/latency first, probabilistic last.
- Monitor
traces_droppedand collector memory. - Test with synthetic traces before production rollout.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.