Helm Release Management and External State
The `--atomic` flag in Helm automate recovery by triggering a rollback if a release fails to reach a ready state within the specified `--timeout` period. This ensures the Kubernetes cluster returns to the last known stable revision of the chart and values.
However, Helm manages the state of Kubernetes objects but does not have native visibility into external state changes, such as database schema migrations executed via init containers or Job hooks during the upgrade process.
When an atomic rollback occurs, the deployment reverts to the previous image and configuration, but the external data modifications remain applied. This creates a potential mismatch between the deployed application version and the current state of the backend data.
How Helm handles the lifecycle of Job hooks during an automatic rollback?
Helm does not automatically "undo" the effects of a Job hook. If a Job with a pre-upgrade or post-upgrade hook completes successfully but the subsequent deployment fails, Helm triggers a rollback to the previous revision. However, the Job itself has already executed. During the rollback:
- Jobs created during the failed upgrade attempt are not automatically deleted by Helm unless specifically configured via the
helm.sh/hook-delete-policy annotation.
- If the hook was a
pre-upgrade job and it succeeded, the rollback reverts the application pods but the data remains.
- If the hook was a
post-upgrade job and failed, the atomic rollback triggers, but the data state may be partially mutated.
Is there a documented mechanism for a compensating transaction?
There is no native, automated mechanism in Helm to trigger a data-level rollback or compensating transaction. Helm operates strictly at the Kubernetes API level. To handle external state, you must implement one of the following patterns:
- Compensating Hooks: Define a
post-rollback hook designed to run a reversal script (e.g., a database migration down). Note that this only runs if the rollback itself is triggered.
- Idempotent Migrations: Ensure your application and database schema changes are backward-compatible so that the older application version can still function against the newer schema.
- Health Checks: Use a
pre-upgrade hook to validate state compatibility before the upgrade begins, preventing the failure-state entirely.
To provide a more specific recommendation, are your database migrations handled via Kubernetes Job hooks or from within the application's init containers?