Fine-tuning Datadog APM Sampling Rate to Balance Cost and Coverage
Sampling rate controls how many transactions Datadog traces, directly affecting APM costs and your ability to spot rare bugs. This post walks through a practical engineering decision: adjusting the rate at runtime, verifying the change, and avoiding common blind spots.
11 Jun 2026, 09:13 UTC

Problem: trace coverage vs APM cost
Every transaction Datadog traces consumes storage and processing capacity. When a team onboarded a new high-throughput service, the default 10% sampling rate meant thousands of traces per hour, pushing monthly APM spend higher than expected. At the same time, reducing the rate risked hiding latency spikes and intermittent errors that only appear in full traces. The engineering question wasnt whether to sample, but how to dial the rate to match the services risk profile and budget constraints.
What the sampling rate actually controls
Datadogs sampling rate is the percentage of incoming requests that get fully traced and stored in APM. A rate of 10% means roughly one in ten requests appears in the transaction list; the rest are dropped after a lightweight span is recorded. The rate appears as a small badge next to each transaction in the APM UI, and you can filter the list by sample rate, service, or error status. This badge is the fastest way to confirm a change took effect without leaving the dashboard.
Adjusting the rate at runtime
Most Datadog language agents support dynamic sampling adjustments. In Ruby, for example, adding Datadog.configure do |config| config.service = my-service config.sampling_rate = 0.05 end and deploying the config file, or calling the method at runtime, updates the rate without a process restart. Python, Node, and Go agents follow similar patterns via the Datadog.configure API or environment variables such as DD_TRACE_SAMPLE_RATE. Changes typically propagate within seconds on the agent side, and the transaction list badge updates accordingly.
# Example: set 5% sampling for a Python service via environment export DD_TRACE_SAMPLE_RATE=0.05 # Or programmatically: from datadog import trace trace.configure(sampling_rate=0.05)
Not all agents honor runtime changes. Java agents, for instance, often require a JVM restart or a new deployment to apply a new sampling rate. If your pipeline includes automated canary releases, a runtime adjustment can serve as an A/B test: keep the default 10% for the control group and lower the rate for the experimental group to evaluate cost impact.
Trade-offs and limitations
- Rare bugs get hidden. Aggressively lowering the sampling rate (e.g., below 1%) may mean a critical error never appears in APM unless you temporarily revert to 100% tracing.
- Agent heterogeneity. Mixed-language services may need different configuration paths; some require restarts, others accept environment variables.
- Cost vs visibility sweet spot. For many services, 5-10% provides sufficient latency visibility while cutting trace volume by half or more.
Verifying the change in the UI
- Open the APM transaction list for the target service.
- Look for the sample-rate badge on individual rows; it reflects the currently active rate.
- Use the UI filter Sample rate to narrow the list and confirm the transaction count drops or rises as expected.
- Navigate to the service details page and check the Sampling rate setting matches the intended value for the environment.
After adjusting the rate, wait a minute or two for agents to push the new configuration, then recheck the badge. If the count does not shift, verify that the agent version supports dynamic sampling and that no conflicting environment variable or config file is overriding the value.
Closing: a practical next step
Pick one low-risk service and reduce its sampling rate from the default 10% to 5%. Monitor the transaction count and APM cost dashboard for a day, then compare error detection rates. If the lower rate meets your coverage needs, roll the change outward. If rare bugs surface, increment the rate back or enable targeted 100% tracing for that services error paths. The sampling rate is a dial not a one-time decision treat it as an ongoing observability tuning knob.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.