Configuring Adaptive Sampling in Jaeger for Production Services
Configure Jaeger 1.x adaptive sampling to hold a steady traces-per-second target across hot and cold services, with verification steps, fallback behavior, and rollback.
17 Aug 2026, 19:30 UTC

The problem: fixed-rate sampling breaks at both ends of your traffic curve
A probabilistic sampler set to 1% works fine at steady load, but it drowns your storage during a traffic spike and starves you of traces from low-volume services that barely see one request per minute. Jaeger's adaptive sampling addresses this by letting the collector compute per-service, per-operation sampling probabilities aimed at a target number of traces per second, then distributing those probabilities to clients through Jaeger's remote sampling mechanism. The useful takeaway: you get roughly uniform trace coverage across hot and cold endpoints without hand-tuning a sampling file every time traffic patterns shift.
This guide assumes Jaeger 1.x (the behavior described applies to the 1.x collector; the Jaeger v2 architecture, built on the OpenTelemetry Collector, uses different configuration keys and should be checked against its own docs before applying anything here).
Prerequisites
- A running Jaeger 1.x collector you can reconfigure and restart.
- A supported storage backend for the adaptive sampling store (Cassandra or Elasticsearch are the commonly documented options; in-memory is not suitable for production because calculated probabilities are lost on restart).
- Clients that poll the collector for sampling strategies: Jaeger client libraries do this natively. OpenTelemetry SDKs do not poll Jaeger's remote sampling endpoint by default, so for OTel-instrumented services you typically apply sampling at the collector or via the OTel SDK's parent-based sampler that respects incoming decisions.
- Prometheus (or equivalent) scraping the collector's metrics endpoint, so you can verify the sampler is actually converging.
How adaptive sampling decides
The collector periodically aggregates the span throughput it observes per service and operation, compares it against a target traces-per-second value, and writes updated probabilities into the sampling store. Clients fetch these strategies (by default over the collector's sampling endpoint) and apply them locally. Because the calculation is based on observed traffic, there is a lag: expect probabilities to settle over minutes, not milliseconds, after a traffic change.
Enabling it on the collector
Adaptive sampling is enabled with collector flags (or their environment-variable equivalents). Run these wherever the collector process is managed — systemd unit, container spec, or Helm values. You need permission to modify and restart the collector deployment.
jaeger-collector \
--sampling.store.type=cassandra \
--sampling.target-samples-per-second=1.0 \
--sampling.delta-tolerance=0.3 \
--sampling.buckets-for-calculation=10 \
--sampling.calculation-interval=1mMeaning of the key knobs:
--sampling.store.type: where calculated probabilities persist. Must match a backend your collector already talks to.--sampling.target-samples-per-second: traces per second the sampler aims for per operation. Start at 1.0 and lower it for very high-volume services.--sampling.delta-tolerance: how far observed throughput can drift from the target before probabilities are recalculated. Higher values mean less churn, slower reaction.--sampling.calculation-interval: how often the leader collector recalculates probabilities.
Only one collector instance performs the calculation (leader election is handled internally when multiple replicas share the store), so this is safe to run with a replicated collector.
Fallback behavior to know about
If a client cannot reach the collector's sampling endpoint, it falls back to whatever default sampler it was constructed with — often a fixed probabilistic rate. During a collector outage this can silently change your trace volume in either direction. Keep an initial, conservative default sampler configured in your applications so the fallback is safe.
Verifying it works
Three checks, in order:
- Strategy endpoint responds. From a host with network access to the collector, query
curl http://<collector-host>:14268/api/sampling?service=<your-service>. You should get a JSON strategy with per-operation probabilistic rates rather than an error or an empty default. Port 14268 is the default HTTP port in Jaeger 1.x; adjust if you remapped it. - Metrics converge. Scrape the collector metrics and watch throughput versus target. Under steady load, stored traces per second for a given operation should hover near
target-samples-per-second, not scale linearly with request volume. - Burst test. Generate a short, large traffic spike against one endpoint. Stored trace count should rise only modestly; the sampler should push the probability down within a few calculation intervals. If stored traces scale 1:1 with the spike, the clients are not picking up remote strategies — check that they are configured to poll the collector and that no SDK-side sampler is overriding the decision.
Operational caveats
- Collector CPU. Strategy calculation and serving add measurable CPU load. Watch collector CPU after enabling, especially with many services and operations.
- Changing the target. Updating
--sampling.target-samples-per-secondrequires a collector restart (or rolling redeploy); the flag is read at startup. Plan the change like any other rollout. - Tail-based needs. Adaptive sampling is head-based: the decision is made at trace start. If you need to keep all error traces regardless of rate, you need tail sampling (available in the OpenTelemetry Collector) rather than Jaeger's adaptive sampler.
- Cold-start gap. A brand-new service has no observed throughput, so it gets an initial probability until the first calculation interval completes. Verify new services emit traces at all before assuming the sampler is broken.
Recovery
Because enabling adaptive sampling changes collector state (the sampling store contents and served strategies), keep a rollback path: redeploy the collector with the previous flags, which reverts to serving the static strategies file or fixed sampler you had before. Stale calculated probabilities in the store are harmless once adaptive sampling is disabled, but you can purge the sampling keyspace/bucket if you want a clean slate. After rollback, re-run the strategy-endpoint check above to confirm clients receive the old configuration.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.