Direct answer to the three questions
Operational risks and benefits of enabling containerd GC in Talos
Benefits. Enabling containerd garbage collection and kubelet image GC provides predictable reclamation of unused layers in /var/lib/containerd on the writable partition. It reduces steady growth from dangling snapshots and unreferenced images left by frequent pulls, rollouts and CI workloads. With thresholds tuned to node disk size, it prevents disk pressure evictions and avoids manual pruning.
Risks. Talos uses an immutable root filesystem. Images and container state live under /var/lib/containerd on the writable partition. Aggressive GC can remove layers that are not currently referenced by running pods but are needed for fast restart, node upgrade, or rollback. Removing images referenced by system pods or Talos-managed workloads can cause pod failures and node instability. Changes to containerd or kubelet settings in Talos require a machine config patch and a controlled reboot, which impacts node availability.
Effect on node upgrades and rollback to previous machine configurations
Node upgrades in Talos are driven by machine config changes and the Talos control plane. The ability to roll back a node to a previous machine configuration does not depend on keeping old container images on disk, it depends on the machine config history and the ability to pull required images after reboot.
Enabling GC does not break the Talos upgrade or rollback mechanism itself. The risk is operational: if GC removes an image that a workload needs immediately after a reboot or rollback, the node will need to pull the image from the registry before pods become Ready. That increases startup time and creates a dependency on registry availability and network. Conservative GC settings that keep recently used images and rely on kubelet image GC thresholds limit this risk.
Monitoring to verify GC reclaims space without disrupting workloads
Monitor node disk usage on the writable partition, not the immutable root. Track /var/lib/containerd size and growth rate over days. Correlate with containerd image counts and sizes, and with kubelet eviction events and image GC events.
- Node filesystem metrics: df for /var, node_filesystem_avail_bytes and node_filesystem_used_bytes for the /var mount.
- Containerd image metrics: number of images, layers, and total size; reference counts for images used by running pods.
- Kubelet metrics: image_fs_available_bytes, image_fs_inodes_free, kubelet_eviction_* counters, and kubelet image GC events.
- Workload health: pod restart counts, image pull latency after node reboot, and time to Ready after upgrades.
Alert on sustained growth of /var/lib/containerd despite GC being enabled, and on image pull failures after reboots.
Likely explanation vs confirmed facts
Confirmed facts
- Talos Linux is an immutable Kubernetes node OS. The root filesystem is immutable; writable state is on /var.
- Containerd stores images and container state under /var/lib/containerd on the writable partition.
- Image cleanup is a combination of containerd internal GC for unused layers and kubelet image GC driven by node disk pressure thresholds configured in the Talos machine config.
- Changes to containerd or kubelet settings require a machine config patch and a controlled reboot to apply.
Likely explanation for steady disk consumption
Steady growth is commonly caused by accumulation of unused image layers and dangling snapshots after frequent pulls, rollouts, or CI workloads. Containerd GC is conservative by default and kubelet image GC only activates when disk usage crosses configured eviction thresholds, so /var can grow substantially before automatic pruning occurs. The brief notes containerd GC policy is disabled by default in this environment, leaving cleanup to manual intervention and causing inconsistent disk usage across clusters. This is plausible but requires current verification against the applied machine config for your Talos version.
Steps needed for this case
- Confirm the full filesystem and top consumers.
df -h
du -sh /var/lib/containerd/* | sort -h
- List containerd images and sizes to quantify unused layers.
ctr -n k8s.io images list
ctr -n k8s.io images ls --size
Correlate with running pods: crictl ps -a and image references.
- Inspect the applied Talos machine config for containerd config and kubelet eviction/imageGC settings.
talosctl machineconfig show
Check containerd config and kubelet config sections for imageGarbageCollectionPolicy, imageGCHighThresholdPercent, imageGCLowThresholdPercent, and eviction hard thresholds.
- Prune unreferenced images safely, keeping images used by running pods and system workloads.
ctr -n k8s.io garbage collect --force
Do not force-remove images referenced by running pods.
- Re-evaluate thresholds to match node disk size and workload churn, then apply via machine config patch and controlled reboot. Monitor /var usage over several days to confirm growth rate changes.
One missing diagnostic detail that changes the recommendation: the current Talos version and the applied machine config values for containerd garbage collection and kubelet imageGcPolicy/eviction thresholds. Without those, threshold tuning and the decision to enable containerd GC cannot be safely scoped.