Controlling Trace Volume with Jaeger’s Remote Sampling: A Pragmatic Guide
High‑traffic services can flood Jaeger with spans. Remote sampling lets you set global, per‑service, and per‑operation rates centrally—without redeploying. Learn how to configure, verify, and trade‑off head‑based sampling versus tail‑based guarantees.
09 Sept 2025, 18:36 UTC

When Trace Volume Becomes a Problem
Once a microservice starts emitting OpenTelemetry spans to Jaeger, the number of spans can grow linearly with traffic. A 1 % sampling rate on a 10 kops service still results in 100 traces per second, which can quickly exhaust storage and slow down queries. The root of the problem is the decision made at the start of a trace: which traces to keep and which to drop. Jaeger solves this with a remote sampling API that lets you push sampling strategies to all instrumented SDKs.
What Remote Sampling Is and Why It Matters
Remote sampling is a head‑based strategy: the SDK pulls a JSON strategy from Jaeger before it creates a trace. The strategy can be a simple probability (e.g., keep 1 % of traces) or a rate limit (e.g., keep at most 5 traces per second). Because the strategy lives in Jaeger, you can change it without redeploying services, making it a pragmatic first line of defense against runaway trace volume.
- Probabilistic sampling keeps a random percentage of traces. It’s simple and works well when traffic is relatively uniform.
- Rate‑limiting sampling caps the number of traces per second, useful for expensive or high‑latency operations that you want to keep in bursts.
- Both can be overridden per service or operation via the strategy JSON.
Configuring Remote Sampling in Jaeger 2.x (Collector)
In Jaeger v2, the Collector is the component that serves the sampling strategy. The configuration is part of the Collector’s YAML file (or command‑line flags) and looks like this:
sampling:
default: "probabilistic(0.01)" # 1 % global rate
rules:
- service: checkout-service
strategy: "probabilistic(0.10)" # 10 % for checkout
- service: debug-service
strategy: "rateLimiting(5)" # 5 traces/sec for debug
When an OpenTelemetry SDK starts a trace, it queries http://collector-host:14267/sampling?service=service-name. The Collector returns a JSON strategy that the SDK applies immediately.
Example Strategy JSON
Below is the exact JSON that the Collector would return for checkout-service:
{
"type": "probabilistic",
"samplingRate": 0.10
}
For debug-service you would see:
{
"type": "rateLimiting",
"maxTracesPerSecond": 5
}
Verifying the Strategy
- Run the Collector locally:
jaeger-collector --config-file=collector.yaml. - Send a trace from a demo app with service name
checkout-service.- Use
curlto query the endpoint:curl -s http://localhost:14267/sampling?service=checkout-service. - Confirm the JSON matches the 10 % strategy above.
- Use
- Open the Jaeger UI (default
http://localhost:16686) and search for traces fromcheckout-service. You should see roughly 10 % of the traces you sent.
Head‑Based vs Tail‑Based Sampling – The Trade‑Off
Head‑based sampling (what Jaeger provides) decides at trace start whether to keep the trace. It’s fast and requires no post‑processing, but it can drop traces that later turn out to be interesting (e.g., those that hit an error or exceed a latency threshold). Tail‑based sampling, which Jaeger does not implement natively, would keep all traces and then filter them after the fact. Tail‑based approaches are usually implemented in an OpenTelemetry Collector pipeline with a tail_sampling processor.
Practical consequence: if your service has a rare error that you want to capture, a head‑based 1 % strategy might miss it. In that case, you can either raise the probability for that operation or add a rate‑limit that ensures you see at least a few traces per second. If you need guaranteed visibility, consider a two‑step pipeline: first sample broadly with Jaeger, then run a Collector tail‑sampler on the exported data.
Limitations to Keep in Mind
- Remote sampling only works if trace context propagates correctly between services. If a downstream service starts a new trace instead of continuing the parent, it will make its own head‑based decision.
- Not all OpenTelemetry SDKs support Jaeger’s remote sampler out of the box. Some require a dedicated
jaeger_remote_samplerplugin or manual configuration. - Sampling decisions are irrevocable once a trace is started. A 0 % probability means the trace never reaches Jaeger, so any instrumentation that depends on the trace (e.g., logging correlation) will be lost.
- Storage backends (Memory/Badger vs Cassandra/Elasticsearch) do not affect sampling; they only determine how many traces are kept after the decision.
Practical Checklist for Production
- Deploy the Collector with a sensible default (e.g., 1 % probabilistic).
- Identify high‑value services or operations and raise their sampling rates.
- For expensive or debug‑heavy endpoints, apply a rate‑limit to keep the volume manageable.
- Periodically query the
/samplingendpoint to ensure the strategy is being served correctly. - Monitor the Jaeger UI for missing traces; if you notice a spike in dropped errors, consider a tail‑sampling pipeline.
- Document the strategy in your observability playbook so that new teams can adjust rates without code changes.
Conclusion
Jaeger’s remote sampling feature gives teams a lightweight, centrally managed way to control trace volume. By combining probabilistic defaults with per‑service overrides and rate‑limits, you can keep costs under control while still capturing the most important traces. Remember that head‑based sampling cannot guarantee retention of error or latency‑critical traces; if that level of assurance is required, augment your pipeline with a tail‑sampler in the Collector.
Start with a conservative global rate, adjust based on operational needs, and keep an eye on the /sampling endpoint to verify that your strategy is being applied as intended.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.