Centralized Trace Sampling in Jaeger Using Remote Strategies via the Agent
Jaeger clients fetch probabilistic or rate‑limiting sampling strategies from the collector through the local agent, enabling dynamic trace‑volume control without code changes or restarts.
01 Oct 2025, 22:04 UTC

Useful answer
Jaeger clients can obtain sampling strategies from the collector through the local agent, letting you adjust trace volume without redeploying code.
How remote sampling works
When a Jaeger‑instrumented service starts, its client library creates a default sampler (often probabilistic with a very low rate, e.g., 0.001). The client then polls the Jaeger agent that runs on the same node—typically via UDP on port 6831 or HTTP on port 14268. The agent forwards the request to the Jaeger collector over gRPC (usually collector:14250). The collector looks up a strategy for the service name (and optionally operation) in its configuration and returns a JSON document such as:
{"strategyType":"probabilistic","probabilisticSampling":{"samplingRate":0.1}}
The client replaces its local sampler with the received strategy and applies it to all newly created spans. Spans that are sampled out are dropped before they leave the host, so no unnecessary network traffic is generated. The polling interval is controlled by the client library (commonly 10–60 seconds); a new strategy takes effect on the next poll, without a restart.
Worked configuration
The collector defines strategies in its YAML configuration. Below is an example that sets a higher probabilistic rate for the payment service, a rate‑limiting cap for its checkout endpoint, and disables tracing for the catalog service.
sampling:
default_strategy:
strategy_type: probabilistic
probabilistic_sampling:
sampling_rate: 0.001
strategies:
- service: payment-service
operation: "*"
strategy_type: probabilistic
probabilistic_sampling:
sampling_rate: 0.1
- service: payment-service
operation: "POST /checkout"
strategy_type: rate_limiting
rate_limiting_sampling:
max_traces_per_second: 5
- service: catalog-service
strategy_type: const
const_sampling:
decision: false
The agent is started as a sidecar or DaemonSet with flags that point it at the collector:
--reporter.grpc.host-port=collector:14250
--reporter.log-spans=false
A Java client (or any language client) is configured to use the remote sampler and to contact the agent:
serviceName=payment-service
sampler.type=remote
sampler.param=1 # ignored when remote is used; kept for compatibility
sampler.remote-agent-address=localhost:6831
On start‑up the client uses the default probabilistic 0.001 rate. After the first successful poll (usually within 10–60 seconds) it receives the strategy for payment‑service and switches to a 10 % probabilistic sampler. Traces for POST /checkout are further limited to five per second by the rate‑limiting sampler.
Verification steps
- Confirm the collector is serving the strategy. From a host that can reach the collector, run:
curl http://collector:14269/sampling?service=payment-service
You should see a JSON document matching the YAML entry (e.g., strategyType = probabilistic, samplingRate = 0.1). Note that exposing this endpoint outside the cluster can leak service names.
- Check collector logs for lines like “Serving sampling strategy for service=payment‑service”.
- Enable DEBUG logging on the client and look for “Fetched sampling strategy” entries that show the received strategy type and parameters.
- Send a test trace (e.g., via a curl to your instrumented endpoint) and open the Jaeger UI. In the trace details, the tags sampler.type and sampler.param should reflect the remote strategy (probabilistic and 0.1, or rate_limiting with max_traces_per_second).
- To test fallback, stop the agent or block network to the collector. After the client’s polling interval expires, it will revert to the default sampler. Verify in the UI that the sampler.type returns to probabilistic and sampler.param shows 0.001 (or whatever your default is).
Limits and common mistakes
- Unreachable agent/collector. If the agent cannot contact the collector, the client falls back to its initial sampler. During an outage this can mean a very low sampling rate and the loss of critical traces. Mitigate by running the agent as a DaemonSet with restart policies and monitoring the agent‑to‑collector link.
- Direct‑to‑collector setups bypass the agent. When using the OTLP exporter or configuring the client to send spans straight to the collector, the remote‑sampling endpoint is never consulted. In those cases you must configure sampling inside the OpenTelemetry SDK (e.g., via environment variables OTEL_TRACES_SAMPLER=parentbased_traceidratio).
- Invalid strategy file. A syntax error in the collector’s YAML/JSON causes the collector to return no strategy, leaving clients on their defaults. Validate the file with a YAML linter or start the collector with --validate-config.
- Rate‑limiting is per‑instance. The token bucket is maintained in each client process, not shared across replicas. Bursty traffic can exceed the limit on a busy instance while idle instances stay under the limit, leading to uneven trace distribution. For a true cluster‑wide cap you need an external throttling mechanism.
- Adaptive sampling removed. Adaptive sampling was deprecated and removed in Jaeger 1.30+. Use rateLimiting combined with an autoscaler or a custom policy if you need dynamic adjustment based on load.
- Exact name matching. The service name sent by the client must match exactly the service field in the collector configuration (case‑sensitive). A typo results in the client receiving the default strategy. Operation wildcards (*) are supported, but a misspelled operation string yields the same fallback.
By keeping the sampling decision in the collector and fetching it via the lightweight agent, you gain centralized control that can be adjusted in response to incidents or changing traffic patterns without touching application code or restarting pods.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.