Taming Trace Volume: Implementing Probabilistic Sampling in Jaeger
Stop overloading your storage with redundant traces. Learn how to use Jaeger's probabilistic sampling to maintain system visibility while drastically reducing observability overhead.
25 Dec 2025, 08:59 UTC

The High Cost of Total Visibility
In a distributed system, tracing every single request is an expensive luxury. When your services handle thousands of requests per second, capturing 100% of traces creates a massive amount of data that can overwhelm your storage backend and introduce significant network overhead. The problem isn't just storage; it's the noise—searching through millions of identical successful requests to find the one outlier that caused a latency spike.
The solution is Probabilistic Sampling. Instead of recording everything, you capture a statistically significant fraction of your traffic. This allows you to maintain a representative view of system performance without paying the performance or storage tax of full instrumentation.
How Probabilistic Sampling Works
Probabilistic sampling makes a binary decision—keep or drop—based on a configured percentage. If you set a sampling rate of 0.1, Jaeger will retain approximately 10% of all traces.
A critical technical detail is how Jaeger handles distributed context. A single request often touches ten different services. If Service A decides to sample the request but Service B decides not to, you end up with a broken trace consisting of fragmented spans. To prevent this, Jaeger uses a sampling.priority tag. The first service in the chain makes the sampling decision and propagates that decision to all downstream services. If the root span is sampled, every subsequent span in that specific trace is also sampled, ensuring complete end-to-end visibility for the selected requests.
Dynamic Tuning via the Sampling Endpoint
Static configuration is risky. If you set your sampling rate too low, you might miss a critical production bug. If you set it too high during a traffic surge, you might overload your Jaeger collector.
Jaeger solves this by providing a remote sampling HTTP endpoint. Rather than hardcoding the rate in a config file and restarting your pods, the Jaeger client libraries can poll the collector for the current sampling strategy. This allows you to dial up the visibility during an active incident and dial it back down once the system stabilizes, all without a redeploy.
Example: Configuring and Verifying Sampling
To implement probabilistic sampling, you typically configure the tracer client within your application code. Below is a conceptual configuration for a Jaeger client using a probabilistic strategy.
// Example configuration for a Jaeger Tracer
// This sets the sampling rate to 5% (0.05)
const config = {
sampler_config: {
type: 'probabilistic',
param: 0.05
},
logging: true,
reporter: {
log_spans: true
}
};
Verification Steps:
- Generate Traffic: Send 1,000 requests to your instrumented service.
- Check Jaeger UI: Search for traces within the specific time window. You should see approximately 50 traces (5% of 1,000).
- Inspect Tags: Open a trace and look for the
sampling.prioritytag. It should be set to1for all spans in that trace. - Test Dynamic Change: If the remote sampling endpoint is enabled, send a POST request to the collector's sampling endpoint to change the rate to 0.20. After the client's next poll, you should see the trace volume increase to ~20% in the UI.
Trade-offs and Limitations
While probabilistic sampling saves resources, it introduces a specific blind spot: the long tail problem. Rare errors or extreme latency spikes (p99.9) may occur in the 95% of traces that you are dropping. If an error only happens once every 1,000 requests and your sampling rate is 5%, there is a high probability you will never see the trace for that specific failure.
To mitigate this, consider combining probabilistic sampling with Adaptive Sampling or using a separate error-reporting mechanism that triggers a trace capture regardless of the sampling rate when a 5xx response is detected.
Actionable Summary
To optimize your Jaeger deployment, move away from constant sampling as soon as you hit production scale. Start with a conservative probabilistic rate (e.g., 0.01 to 0.05), enable the remote sampling endpoint for runtime flexibility, and always verify that your sampling.priority tags are propagating correctly across service boundaries to avoid fragmented traces.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.