Answer
Datadog does not publish the exact numeric thresholds that its adaptive sampling algorithm uses to weigh latency, error rate, and throughput. The algorithm internally computes a score from recent span latency (e.g., 95th‑percentile), error percentage, and request volume, then derives a per‑service sample rate that tries to keep overhead low while preserving spikes and errors.
Likely explanation (based on product behavior)
The sampler likely raises the sample rate when:
- Latency exceeds a dynamic high‑latency band (often defined relative to the service’s recent baseline).
- Error rate rises above a low‑percentage threshold (e.g., >0.5%).
- Throughput drops, because low volume services are sampled more heavily to retain statistical significance.
Conversely, when traffic is high and latency/error stay within normal bands, the sample rate is driven down.
Confirmed facts from Datadog
- The effective sample rate for each service is visible in the UI under
APM → Trace Analytics → Sampling and via the /api/v2/trace/sampling endpoint. - Users can override adaptive sampling by setting a fixed rate with the environment variable
DD_TRACE_SAMPLING_RULES or through the UI Sampling Rules page. - Infrastructure metrics (host CPU, memory, etc.) are collected independently of trace sampling, so they remain available for correlation.
Steps to assess and mitigate loss of bottleneck visibility
- Check the current sampling rate for the suspect service in the Sampling UI.
- If the rate is significantly below 1.0 (e.g., <0.2) and you observe latency spikes in host metrics but few traces, the sampler is likely discarding those spikes.
- To verify, temporarily set a fixed sampling rate of 1.0 for that service:
DD_TRACE_SAMPLING_RULES='{"sample_rate":1.0}'
- Redeploy the agent, then re‑examine the trace view for the previously missing high‑latency spans.
- If the spikes appear, decide whether to keep a higher fixed rate or to craft a rule that samples based on latency/error thresholds (e.g.,
DD_TRACE_SAMPLING_RULES='[{"service":"my-service","sample_rate":0.5}]'). - Monitor ingest volume after any change to ensure costs remain acceptable.
Missing diagnostic detail
To give a precise recommendation, I would need to know the current reported sampling rate for the service in question (as shown in the Sampling UI). If you can share that value, I can advise whether adjusting the rule is necessary.