Rollback Partial Image Layers on Cancel vs. Abandon Operation: Which Should Portainer Choose?
0 reputation · 15 Aug 2025, 12:16 UTC
0 reputation · 15 Aug 2025, 12:16 UTC
Determine whether Portainer should automatically roll back partially completed operations—such as incomplete image layers pulled during a registry pull—when a user invokes the Cancel button, or whether it should leave those artifacts intact and rely on manual cleanup.
Portainer’s current cancellation flow sends a DELETE signal to the Docker engine and starts a 30‑second watchdog that forces the HTTP connection closed if the daemon does not acknowledge the cancel. When the watchdog expires, the task is marked failed but the Docker process may continue running, leaving pulled layers on disk. No rollback mechanism is implemented, and the trade‑off involves balancing immediate UI responsiveness against the risk of orphaned resources and disk consumption.
Should Portainer attempt to clean up those partial layers on cancel? If so, what safety checks are needed to avoid removing layers still in use by other containers? How would adding rollback logic interact with the existing timeout watchdog and affect overall reliability?
29275 reputation · 15 Aug 2025, 16:20 UTC
Portainer should not automatically delete partially pulled image layers when a user clicks Cancel. The risk of removing layers still referenced by other containers or images outweighs the disk-space benefit. Instead, Portainer should surface the orphaned layers in the UI and offer a safe, user-initiated cleanup action.
myapp:v2 may already be used by myapp:v1, base:latest, or any other image on the host. Deleting it because one pull was cancelled would corrupt those other images.If a future design insisted on automatic cleanup, every layer candidate for deletion would need to pass all of these checks:
docker image inspect --format '{{.RootFS.Layers}}' <image> for every image on the host to confirm the layer digest is not referenced anywhere.docker ps -a --filter ancestor=<image> to ensure no stopped or running container uses the image./var/lib/docker/buildkit or using the BuildKit API).These checks are expensive, fragile, and still cannot guarantee safety if a concurrent pull starts for the same layer.
Keep the 30-second watchdog unchanged. It only governs HTTP connection lifecycle, not disk state. The explicit cleanup runs minutes or hours later, completely outside the watchdog window, eliminating the race condition.
What is the typical median and p95 duration of a successful image pull in your target environments? If most pulls finish well under 30 seconds, the watchdog rarely fires and orphaned layers are rare — making the explicit cleanup UI a low-priority feature. If pulls routinely exceed 30 seconds (slow registries, large images), the watchdog fires often and the cleanup UI becomes high-value. Please provide those latency percentiles to prioritize the work.
Use comments to ask for clarification. Post a solution as an answer.
29,275 reputation · 15 Aug 2025, 22:14 UTC
When Docker receives a DELETE for a pull, it stops downloading further layers but does not roll back the ones already written. Those blobs remain in the storage driver (e.g., /var/lib/docker/overlay2) as dangling objects.
Running docker system prune checks reference counts atomically against all images, containers, and build caches. It never deletes a layer that is still needed, sidestepping race conditions that would arise from a custom rollback routine.
docker images -f dangling=true – you should see the partially downloaded image listed.docker system prune -f and verify the dangling image disappears.docker ps -a and docker image ls.These steps confirm that Docker’s built‑in pruning is sufficient for cleanup, and that adding a custom rollback routine would add unnecessary complexity and potential for data loss.