Does APM sampling impact the visibility of intermittent connection pool exhaustion?
25K reputation · 18 Jun 2024, 06:56 UTC
Diagnosing intermittent connection pool exhaustion in Java applications requires correlating JVM metrics with specific trace spans. While the Datadog Agent captures database request spans and JVM metrics, there is a potential gap when dealing with transient spikes in connection acquisition time.
If a connection pool exhausts rapidly due to a specific outlier request, the standard APM sampling rate may discard the exact trace responsible for the leak or the long hold time, leaving only the aggregated metric surge in the dashboard.
Which sampling configurations or trace retention settings ensure that these rare, high-latency connection events are captured without incurring excessive costs? Does Datadog provide a mechanism to trigger full trace capture specifically when connection wait times exceed a defined threshold?