Balancing APM Visibility and Cost: Choosing a Datadog Trace Sampling Strategy
Learn how to choose between rate-based and priority-based trace sampling in Datadog APM to reduce ingestion costs without losing critical error data.
07 Feb 2026, 23:42 UTC

The Trace Ingestion Dilemma
High-throughput applications generate millions of traces per minute. Ingesting every single trace into Datadog provides total visibility but often leads to unsustainable costs and "noise" that obscures actual performance regressions. The core challenge is reducing the volume of ingested data without losing the specific traces—errors, timeouts, and p99 latency spikes—that are required for debugging.
The goal is to move from 100% ingestion to a strategy where you only pay for the data that provides actionable insights.
Comparing Sampling Strategies
Datadog APM primarily offers two ways to manage trace volume: Rate-based sampling and Priority-based sampling. The choice depends on whether your priority is strict budget predictability or high-fidelity error capture.
| Feature | Rate-Based Sampling | Priority-Based Sampling |
|---|---|---|
| Mechanism | Fixed percentage of all traces are kept. | Keeps 100% of errors/outliers; samples others. |
| Cost Predictability | High (Linear to traffic). | Variable (Spikes during incidents). |
| Error Visibility | Probabilistic (may miss rare errors). | Deterministic (captures all errors). |
| Configuration | Simple environment variable. | Requires priority rules and agent config. |
Trade-offs and Decision Logic
When to use Rate-Based Sampling
Rate-based sampling is best for stable, high-volume services where you are looking for general performance trends rather than hunting for a "needle in a haystack" bug. Because it samples uniformly, it provides a statistically accurate representation of your average latency, but it is dangerous for services with low error rates; if you sample at 1%, you will only see 1% of your errors.
When to use Priority-Based Sampling
Priority-based sampling is the standard for production environments where reliability is critical. It ensures that any trace marked as an error or exceeding a specific latency threshold is preserved, while the "healthy" traces are sampled aggressively to save costs. The primary risk is ingest spikes: if a service begins failing globally, 100% of those traces will be sent to Datadog, potentially exceeding your budget or hitting ingestion limits.
Implementation Guide
Sampling configurations are typically applied via environment variables in your application container or deployment manifest. These changes require a restart of the instrumented service to take effect.
Option A: Implementing Rate-Based Sampling
To keep only 10% of all traces, set the following environment variable in your application environment:
DD_TRACE_SAMPLE_RATE=0.1
Option B: Implementing Priority-Based Sampling
To enable priority sampling, you must enable the feature and define the sampling rate for the "non-priority" (healthy) traces. To keep all errors but only 1% of healthy traces, use:
DD_TRACE_PRIORITY_SAMPLING_ENABLED=true
DD_TRACE_PRIORITY_SAMPLING_RATE=0.01
Validating the Sampling Result
Since sampling happens at the tracer level, you cannot simply look at the total trace count to verify the configuration. Use these three methods to ensure the strategy is active:
1. Trace Search Filter
Navigate to APM > Trace Search in the Datadog UI. Apply the filter sampling.priority:user. If priority sampling is working, you should see traces that were explicitly kept due to priority rules, even if the overall volume is low.
2. Metric Analysis
Query the following metrics via the Metrics Explorer to compare the ratio of kept traces versus total generated traces:
trace.analytics.kept: The number of traces actually ingested.trace.analytics.priority: The number of traces kept specifically because they were high priority.
3. Agent Log Verification
Check the Datadog Agent logs on the host or container. Search for the string "trace sampling mode" to confirm the agent is recognizing the sampling configuration provided by the instrumented application.
Limitations and Risks
- Tagging Dependency: Priority sampling relies on the tracer correctly identifying errors. If your code catches exceptions but does not mark the span as an error (e.g., via
span.set_tag('error', True)), those traces will be treated as healthy and may be sampled out. - State Change: Because these settings are read at startup, you must perform a rolling restart of your pods or instances to apply new rates.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.