k3s server memory grows steadily: separating a real leak from cache and containerd
0 reputation · 24 Feb 2024, 06:42 UTC
k3s ships the control plane, kubelet, containerd, flannel, CoreDNS and ServiceLB as a single binary per node. When the k3s server process shows steady RSS growth on a node with otherwise stable workloads, node-level metrics alone cannot say whether that growth is a genuine leak in one embedded component, page cache accounted to the process, or containerd image and content-store expansion.
k3s does expose Go pprof profiling through the supervisor when debug settings are enabled, so heap snapshots taken at two points under steady load should, in principle, separate control-plane growth from cache. The hard part is interpretation. Historical leak reports have spanned subsystems such as the supervisor's tunnel (remotedialer) connections and metrics handling, and whether any particular leak is fixed depends on the exact release line — release notes have to be checked per version, and profiling flags and memory baselines also shift between releases.
Before escalating a growth pattern as a bug, what evidence is actually sufficient — a diffed heap profile from the supervisor's pprof endpoint, RSS trended alongside node allocatable memory and containerd store size, or both? And once true process growth is confirmed, what is the dependable way to attribute it to a specific embedded component and match it against closed issues in the relevant k3s release notes?