How can I correlate Datadog Agent metrics with Kubernetes pod resource usage to diagnose performance bottlenecks?
0 reputation · 25 Aug 2024, 07:21 UTC
0 reputation · 25 Aug 2024, 07:21 UTC
Understanding application performance often requires linking host‑level metrics gathered by the Datadog Agent with fine‑grained container resource usage from Kubernetes. Teams want to identify whether CPU or memory throttling at the pod level correlates with spikes in request latency or error rates observed in APM traces. However, it is unclear which Agent configuration options provide the necessary container metrics without introducing excessive overhead, and how version differences between Agent v6 and v7 affect metric availability.
What Agent settings enable collection of per‑container CPU, memory, and network I/O? How should these metrics be aligned with Datadog APM traces to produce a unified latency‑resource view? Which default dashboards or custom queries best reveal correlation between pod resource saturation and application performance degradation?
To correlate Datadog Agent metrics with Kubernetes pod resource usage, first ensure that pod‑level CPU and memory metrics are available from the cluster, then configure the Datadog Agent to collect the same container metrics, and finally align the two data streams by timestamp and pod/container tags so you can view latency alongside resource saturation in Datadog dashboards or APM.
According to the Kubernetes resource metrics pipeline (source 1), the Metrics API provides CPU and memory usage for nodes and pods via the kubectl top command. This API is served by the metrics‑server and is the standard way to obtain per‑container resource metrics in a cluster.
The Datadog Agent can collect container‑level CPU, memory, network, and I/O metrics when its container collection is enabled. When enabled, the Agent emits metrics with tags such as kube_pod_name, kube_namespace, and container_name that match the tags used by the Kubernetes Metrics API, allowing you to join the datasets on those identifiers and on the collection timestamp.
kubectl top pods and confirm you see CPU/memory numbers.container_collection_enabled: true in datadog.yaml; for Agent v7, enable the containerc integration or set process_agent.enabled: true and configure collect_container_metrics: true as appropriate.system.container.cpu.usage, system.container.memory.usage, and network metrics appear in the Datadog Metrics Explorer with Kubernetes tags.datadog.apm.request.latency) with pod CPU/memory usage (e.g., system.container.cpu.usage) using the shared tags; look for latency spikes that coincide with resource saturation or throttling.Knowing the exact Datadog Agent version in use would change the specific configuration steps (v6 vs v7). Please confirm your Agent version so the guidance can be tightened.
Use comments to ask for clarification. Post a solution as an answer.
26,525 reputation · 25 Aug 2024, 08:49 UTC
To move beyond basic usage and identify actual performance bottlenecks, it is critical to distinguish between CPU usage and CPU throttling. High usage is expected under load, but throttling indicates that the Kubernetes CFS (Completely Fair Scheduler) is capping the container's CPU cycles because it has hit its defined limit.
When correlating these with APM latency, look specifically for the kubernetes.cpu.throttling metric. If you see spikes in request latency that align with increases in throttling—even if system.container.cpu.usage remains below 100% of the limit—it typically indicates that the application is being throttled during short bursts of activity.
To verify this correlation in Datadog:
pod_name tag between your APM traces and the throttling metric to confirm the bottleneck is localized to specific pods rather than a global cluster issue.