Built‑in /metrics endpoint versus external Prometheus for Rancher memory leak detection under a 15‑second scrape interval
0 reputation · 15 May 2020, 11:17 UTC
0 reputation · 15 May 2020, 11:17 UTC
Goal: Choose between Rancher’s native /metrics endpoint and an external Prometheus stack from the monitoring catalog to identify upward memory trends in the server or agent components.
Constraint: The Rancher HA setup recommends a metrics scrape interval of no more than 15 seconds to limit overhead, yet the agent may be restarted after an OOM kill, which could affect metric continuity.
Uncertainty: It is unclear whether the built‑in endpoint provides sufficient granularity for leak detection under the 15‑second interval, or whether the added complexity of an external Prometheus deployment is justified given possible agent restarts.
Which approach offers adequate sensitivity to detect a memory leak while respecting the scrape‑interval limit? How does agent restart after an OOM kill impact the reliability of each method? Is the additional resource consumption of an external Prometheus stack acceptable for leak detection in a production HA Rancher deployment?
29275 reputation · 15 May 2020, 13:47 UTC
Use an external Prometheus scraping Rancher's /metrics endpoint. The built-in endpoint is the telemetry source, not a detection method: it exposes point-in-time samples with no retention, no baseline, and no alerting. Leak detection requires comparing memory over hours, and only Prometheus (or an equivalent TSDB) gives you that. This is not an either/or choice — Prometheus consumes the built-in endpoint.
A memory leak is a trend, not a reading. Confirming one requires showing that resident memory or the post-GC heap floor keeps rising while workload is flat. A single scrape of /metrics cannot distinguish a leak from cache growth, TLS/session pools, or a one-off migration spike. At a 15-second interval you will also miss short allocation bursts and GC sawtooth within individual samples — which is fine, because the signal you care about is the rising floor across scrapes, judged over 24–72 hours with rate()/increase() and post-GC minimums, not any single high value.
An OOM kill resets the process's memory counters to near zero, which breaks continuity for both approaches equally — but only Prometheus lets you see the pattern. With restarts visible (via kube_pod_container_status_restarts_count or changes() on start time), a sawtooth of rising memory followed by an OOM reset, repeated across cycles, is itself strong leak evidence. With only the raw endpoint, each restart silently erases the trend and you never accumulate proof. One caveat: a 15s scrape can miss the final pre-OOM peak, so correlate with container OOM events rather than expecting the last scrape to show the true maximum.
Scrape the Rancher management server and the relevant agents separately — a leak in rancher, cattle-cluster-agent, fleet, or a controller has different blast radius and remediation. Useful series, where exposed:
process_resident_memory_bytes and container_memory_working_set_bytesgo_memstats_heap_inuse_bytes / go_memstats_alloc_bytes (post-GC floor)Alert only when memory grows over a sustained window and remains elevated after GC — not on single-scrape spikes.
Generally yes for a production HA deployment, with one real risk: cardinality. Rancher-managed clusters can generate high-cardinality labels, and an unbounded Prometheus can itself become the memory problem. Keep the 15s interval (adequate for leak triage), drop or aggregate noisy labels, and set retention appropriate to multi-day trend analysis. The cost is modest compared to an undiagnosed leak taking down the management plane.
/metrics endpoint once (port-forward or via the service) to confirm reachability, auth requirements, and which Go/process metrics actually exist — this varies by Rancher version and distribution.One open detail that would sharpen this: which component you suspect (server vs cluster agent), since that changes which metrics and remediation apply.
Use comments to ask for clarification. Post a solution as an answer.
No question comments on this page.