Balancing Trace Fidelity and Costs with Jaeger’s Sampling Strategies
Learn how to use Jaeger’s probabilistic and rate‑limiting sampling to keep observability high while controlling storage, and see a concrete example that shows how to configure and verify the settings.
25 Aug 2026, 01:17 UTC

Problem: Too Many Traces, Too Few Resources
In a busy microservice environment, every request can generate dozens of spans. If every span is shipped to Jaeger, the collector swells, the backend indexes fill, and query latency rises. Yet turning sampling off leaves you with no observability during incidents. The engineering decision is clear: how do you keep enough trace fidelity to debug while keeping storage and network usage under control?
Thesis: Use Jaeger’s Dual Sampling Strategies Wisely
Jaeger offers two complementary sampling modes—probabilistic and rate‑limiting—that can be applied at the agent, collector, or service level. By configuring them together, you can preserve a statistically representative sample of traces and guarantee a hard cap on the number of traces that reach the backend during traffic spikes.
Section 1: Probabilistic Sampling – Random but Representative
Probabilistic sampling selects each trace with a fixed probability p. If you set p to 0.01, roughly 1 in 100 traces will be sent. This strategy is unbiased: the sample distribution of latency, error rates, and other metrics mirrors the full traffic, assuming the probability is truly random.
Configuration example (agent jaeger-agent.yaml):
agent:
sampling:
type: probabilistic
param: 0.01 # 1% sampling
When a trace is sampled, all its spans are retained. If the trace is not sampled, no spans are sent. The agent reports the current sampling rate via the jaeger_agent_sampling_rate metric.
Verification Steps
- Run
curl http://localhost:14245/metrics | grep jaeger_agent_sampling_rateand confirm the value is 0.01. - Generate a synthetic workload that creates 10,000 requests. After a minute, query Jaeger’s
/api/tracesendpoint for trace count. You should see ~100 traces, within statistical variance. - Inspect a trace in the UI. All spans should be present; no missing intermediate spans.
Section 2: Rate‑Limiting Sampling – Caps the Volume
Rate‑limiting sampling ensures that no more than N traces per second reach the collector. It is useful when traffic suddenly spikes and you want to avoid a burst of traces that could overload storage.
Configuration example (collector jaeger-collector.yaml):
collector:
sampling:
type: ratelimit
param: 500 # 500 traces per second
When the incoming trace rate exceeds 500/s, the collector drops excess traces. This can lead to missing some traces during peaks, but the trade‑off is predictable storage usage.
Verification Steps
- Expose the collector metrics endpoint and watch
jaeger_collector_traces_in_totalandjaeger_collector_traces_dropped_totalcounters. - Simulate a traffic burst (e.g., using
wrkorhey) that exceeds 500/s. Verify thattraces_dropped_totalincreases whiletraces_in_totalstays capped. - Check the Jaeger UI for a trace that started during the burst. If it is missing, the rate‑limit worked.
Section 3: Combining Strategies – A Practical Pattern
In many deployments, the agent runs probabilistic sampling to keep a representative set, while the collector enforces a hard cap. This combo ensures that even if the agent’s probability is high during a surge, the collector will throttle the volume.
Example: agent sampling 5% (0.05) but collector limited to 200 traces/s.
# jaeger-agent.yaml
agent:
sampling:
type: probabilistic
param: 0.05
# jaeger-collector.yaml
collector:
sampling:
type: ratelimit
param: 200
Result: You receive up to 200 traces per second, each trace having a 5% chance of being sampled. During periods of low traffic, you’ll still get 5% of traces; during spikes, you’ll get the best 200 traces.
Section 4: Trade‑Offs and Limitations
- Missing Rare Errors – With a low probabilistic rate, a rare error path may never be sampled. To mitigate, use adaptive sampling (OpenTelemetry’s API) that increases rate when an error threshold is crossed.
- Trace Loss During Spikes – Rate‑limiting can drop early or late spans in a trace, breaking end‑to‑end visibility. If you need full traces for critical services, consider a higher per‑service sampling override.
- Duplicate Trace IDs – Ensure each service generates unique trace IDs; otherwise, the backend may merge distinct traces, skewing metrics.
- Configuration Drift – When scaling services, remember to revisit sampling configs; stale rates can cause either data overload or blind spots.
Actionable Closing: How to Decide Your Sampling Settings
- Define your observability goals. If you need 99th‑percentile latency, a 1% probabilistic rate may suffice. If you need to catch every error, increase the rate or use adaptive sampling.
- Measure baseline. Deploy Jaeger with no sampling, capture storage and network usage for a day.
- Apply a conservative probabilistic rate (e.g., 0.01). Verify that
jaeger_agent_sampling_ratematches. - Set a collector rate‑limit that matches your backend capacity. Use metrics to ensure
traces_dropped_totalstays below a small threshold. - Monitor and iterate. Use Jaeger’s
/metricsendpoint and UI percentiles to confirm that sampling hasn’t introduced bias. Adjust rates if you notice missing critical traces.
By balancing probabilistic and rate‑limiting sampling, you can keep Jaeger lightweight while still gaining actionable insights into your system’s behavior.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.