Short answer
Use an external Prometheus scraping Rancher's /metrics endpoint. The built-in endpoint is the telemetry source, not a detection method: it exposes point-in-time samples with no retention, no baseline, and no alerting. Leak detection requires comparing memory over hours, and only Prometheus (or an equivalent TSDB) gives you that. This is not an either/or choice — Prometheus consumes the built-in endpoint.
Why the endpoint alone is insufficient
A memory leak is a trend, not a reading. Confirming one requires showing that resident memory or the post-GC heap floor keeps rising while workload is flat. A single scrape of /metrics cannot distinguish a leak from cache growth, TLS/session pools, or a one-off migration spike. At a 15-second interval you will also miss short allocation bursts and GC sawtooth within individual samples — which is fine, because the signal you care about is the rising floor across scrapes, judged over 24–72 hours with rate()/increase() and post-GC minimums, not any single high value.
Effect of OOM restarts
An OOM kill resets the process's memory counters to near zero, which breaks continuity for both approaches equally — but only Prometheus lets you see the pattern. With restarts visible (via kube_pod_container_status_restarts_count or changes() on start time), a sawtooth of rising memory followed by an OOM reset, repeated across cycles, is itself strong leak evidence. With only the raw endpoint, each restart silently erases the trend and you never accumulate proof. One caveat: a 15s scrape can miss the final pre-OOM peak, so correlate with container OOM events rather than expecting the last scrape to show the true maximum.
What to scrape and graph
Scrape the Rancher management server and the relevant agents separately — a leak in rancher, cattle-cluster-agent, fleet, or a controller has different blast radius and remediation. Useful series, where exposed:
process_resident_memory_bytes and container_memory_working_set_bytesgo_memstats_heap_inuse_bytes / go_memstats_alloc_bytes (post-GC floor)- Restart counts and controller workqueue depth
Alert only when memory grows over a sustained window and remains elevated after GC — not on single-scrape spikes.
Is the Prometheus overhead justified?
Generally yes for a production HA deployment, with one real risk: cardinality. Rancher-managed clusters can generate high-cardinality labels, and an unbounded Prometheus can itself become the memory problem. Keep the 15s interval (adequate for leak triage), drop or aggregate noisy labels, and set retention appropriate to multi-day trend analysis. The cost is modest compared to an undiagnosed leak taking down the management plane.
Verification before you commit
- Query the Rancher
/metrics endpoint once (port-forward or via the service) to confirm reachability, auth requirements, and which Go/process metrics actually exist — this varies by Rancher version and distribution. - Graph RSS/working-set and Go heap for 24–72h at 15s; compare post-GC floors during steady load and after a controlled restart.
- If growth is monotonic, capture a heap/profile snapshot from the suspected component at high memory and compare against a post-restart baseline.
One open detail that would sharpen this: which component you suspect (server vs cluster agent), since that changes which metrics and remediation apply.