Datadog APM Service Map persistence after abrupt process termination
0 reputation · 14 Dec 2022, 22:07 UTC
The Datadog Service Map relies on span metadata to construct and visualize dependency graphs. When a service process crashes abruptly or stops sending traces, the corresponding nodes and edges often persist in the UI for a specific duration.
There is currently uncertainty regarding the exact timeout-defined logic that governs how long these stale nodes remain active before being pruned. In high-throughput environments where local buffering and high-cardinality tags might increase aggregation delays, it becomes difficult to distinguish between a healthy service experiencing latency and a dead service that has ceased functioning.
What is the default duration for which a node remains visible in the Service Map after the final span is received? Is there a user-configurable setting to manually shorten this timeout to ensure more accurate real-time topology representation?