Managing Dynamic Trace Volume with Jaeger Remote Sampling
Learn how to use Jaeger Remote Sampling to dynamically adjust trace rates across services without redeploying code, including configuration examples and verification steps.
07 Jul 2026, 04:06 UTC

Solving the Trace Volume Dilemma
In high-traffic environments, capturing every single request (100% sampling) creates overwhelming storage costs and network overhead. Conversely, a hard-coded sampling rate requires a full service redeployment to change, which is impractical during an active production incident when you need higher visibility.
The solution is Remote Sampling. This mechanism allows you to centrally define sampling strategies in the Jaeger Collector and propagate those rates to your services in real-time without restarting your application pods or servers.
How Remote Sampling Works
Remote sampling shifts the decision-making logic from the application code to the infrastructure. The flow follows a specific polling chain:
- Collector: Acts as the source of truth. It hosts a configuration file defining which services should be sampled and at what rate.
- Agent: Acts as a local proxy. It polls the Collector's admin port and caches the strategies locally.
- Client (SDK): The instrumented application polls the local Agent to determine if it should sample the current request.
Configuration Example
To implement this, you must configure the Collector with a sampling strategy file. Below is a representative YAML configuration for a Jaeger Collector.
# sampling.yaml
sampling:
strategies:
- service: "payment-gateway"
type: "probabilistic"
param: 0.1 # Sample 10% of requests
- service: "auth-service"
type: "ratelimiting"
param: 100 # Sample maximum 100 requests per second
default_strategy:
type: "probabilistic"
param: 0.01 # Sample 1% for all other services
Once the Collector is running with this configuration, the application client must be configured to use the remote sampler. In a Java or Go client, this typically involves setting the sampler type to remote and pointing it to the Agent's sampling port (default 5778).
Operational Constraints and Limitations
While powerful, remote sampling introduces specific architectural dependencies and delays:
- Agent Dependency: Remote sampling requires a Jaeger Agent running as a sidecar or daemon. If your architecture uses an OpenTelemetry Collector to send traces directly to the Jaeger Collector, the remote sampling polling mechanism is bypassed.
- Propagation Latency: Changes are not instantaneous. There is a combined delay consisting of the Agent's poll interval from the Collector and the Client's poll interval from the Agent.
- Granularity: Strategies are applied at the service level. You cannot use remote sampling to target a specific API endpoint or operation; that requires custom sampling logic within the application code.
Common Implementation Pitfalls
| Mistake | Impact | Correction |
|---|---|---|
| Blocking Port 14269 | Agents cannot fetch strategies from Collector. | Ensure firewall/security groups allow traffic on the Collector admin port. |
Missing default_strategy |
Unspecified services may default to an extremely low rate (e.g., 0.001). | Always define a global fallback strategy in the YAML. |
Using const sampler |
The application ignores the Agent's instructions. | Set sampler.type=remote in the client configuration. |
Verifying the Sampling Chain
To ensure your configuration is active, perform these checks in order:
- Verify Collector Exposure: Check if the Collector is serving the strategy. Run this from a machine with access to the Collector:
curl http://<collector-host>:14269/sampling - Verify Agent Cache: Check if the Agent has successfully polled the Collector. Run this from the application host:
curl http://localhost:5778/sampling?service=payment-gateway
The response should return the JSON representation of the strategy defined in your YAML. - Verify Client Adoption: Enable
DEBUGlevel logging for the Jaeger client. Look for log entries statingRemote sampler: using strategy...to confirm the SDK has updated its internal rate.
Rollback Procedure
If a sampling change causes a spike in resource usage or a loss of critical data, revert the sampling.yaml file to the previous version and restart the Collector (or trigger a config reload if supported by your deployment method). The Agents and Clients will automatically converge to the previous state during their next polling cycle.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.