Using Jaeger Adaptive Sampling to Control Trace Volume in Bursty Services
Learn how Jaeger's adaptive sampling automatically tunes trace collection to keep storage costs low while preserving visibility of slow or error requests.
10 May 2026, 11:54 UTC

Problem: Trace volume spikes overwhelm storage
When a service experiences sudden traffic bursts, a fixed‑rate probabilistic sampler can either generate too many traces (filling storage and increasing backend load) or too few (losing visibility of slow or error requests). Teams often resort to manual tuning, which is reactive and error‑prone.
How adaptive sampling works
Jaeger’s adaptive sampling feature creates a feedback loop between the agent and the collector. The agent exports per‑operation statistics (call count, latency, error rate) to the collector. The collector computes a new sampling probability for each operation that aims to keep the observed trace rate near a target maxTracesPerSecond while respecting a per‑operation ceiling maxOperations. The updated probabilities are pushed back to agents, typically within a few seconds, allowing the sampler to react to traffic changes without operator intervention.
Configuration example
To enable adaptive sampling you need a Jaeger version ≥ 1.20 and a storage backend that supports the required metrics (e.g., Elasticsearch, Cassandra). The configuration lives in the collector’s YAML file.
# /etc/jaeger/collector-config.yml sampling: type: probabilistic param: 0.001 # base rate applied when adaptive logic is not yet active adaptive: maxOperations: 2000 # max distinct operations to track maxTracesPerSecond: 500 # target trace ingestion rateAfter editing the file, restart the collector:
# Run on the collector host; requires sudo or the jaeger user sudo systemctl restart jaeger-collectorVerify the change:
- Check the collector logs for lines containing
adaptive-samplingand the configured thresholds. - In the Jaeger UI, navigate to the Sampling tab; you should see per‑operation probabilities that adjust as you generate traffic.
- Optionally, use the Jaeger CLI to query the strategy:
jaeger query --lookup-sampling-strategy(replace with your endpoint).
Trade‑offs and limitations
Adaptive sampling reduces storage pressure but introduces a few operational considerations:
- Version and storage requirement: Older Jaeger releases (< 1.20) lack the feature; some storage backends do not export the necessary metrics, limiting where it can be deployed.
- Configuration sensitivity: Setting
maxTracesPerSecondtoo high can cause over‑sampling during spikes, increasing resource usage; setting it too low may under‑sample and hide critical latency outliers. - Feedback latency: While the loop reacts within seconds, very rapid bursts (sub‑second) may still see a brief mismatch between actual traffic and applied probability.
Practical way to validate the result: after a load test (e.g., using hey -z 1m -c 50 http://your-service), observe the sampling probability in the UI. If it rises during the test and falls afterward, the adaptive loop is functioning as intended.
Actionable closing
If you operate a Jaeger deployment with variable traffic, try enabling adaptive sampling:
- Confirm you are on Jaeger 1.20+ and your storage backend supports the metrics.
- Add the
sampling.adaptiveblock to your collector configuration, start with a modestmaxTracesPerSecond(e.g., 20% of your peak observed trace rate). - Restart the collector and monitor the UI and logs for adaptive‑sampling messages.
- Tune
maxTracesPerSecondandmaxOperationsbased on observed trace volume and latency goals. - Document the chosen values and set up an alert if the collector logs show repeated
adaptive-samplingwarnings, which may indicate mis‑configuration.
By letting the system self‑adjust, you retain visibility of important traces while keeping storage costs under control.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.