Atomic install rollback guarantees under API server contention
0 reputation · 05 Nov 2023, 18:57 UTC
Context
Helm's --atomic flag promises that a failed install rolls back all created resources, leaving no partial state for a retry to duplicate. The documentation states that hook phases are respected and release names must be unique per cluster.
Uncertainty
The guarantee depends on successful API server operations. Network partitions, API throttling, or admission-webhook timeouts can interrupt the rollback itself, potentially leaving resources in an indeterminate state. A subsequent helm install --atomic with the same release name may then encounter conflicts or silently overwrite resources created by the aborted attempt.
Goal
Determine the practical boundaries of the atomic guarantee when the control plane is under stress, so that retry logic can be designed without assuming perfect rollback.
- Does a failed atomic install always leave the release in a
failedorpending-installstatus, or can it become stuck in an intermediate state? - What verification steps confirm that no orphaned resources remain after a rollback triggered by API server errors?
- Are there documented patterns for safe retries when the cluster exhibits transient API unavailability?