To ensure zero-downtime during database schema migrations in Helm, you must use the pre-upgrade and pre-install hook strategies combined with a backward-compatible schema design.
The challenge is not just the hook itself, but the overlap period. When Helm executes a pre-upgrade hook, the migration job completes before the Deployment begins its RollingUpdate. If the migration removes or renames columns that the currently running application pods depend on, those pods will crash before the new pods are even ready.
The Zero-Downtime Workflow
To prevent service interruption, the migration must follow a multi-phase compatibility pattern:
- Phase 1: Additive: Use a
pre-upgrade hook to perform additive changes. This includes adding new columns, tables, or indexes while leaving existing structures untouched.
- Phase 2: Deploy: Helm proceeds with the
RollingUpdate. The new pods use the new schema features, while the old pods continue using the old features (which still exist).
- Phase 3: Cleanup: Once the deployment is fully successful, a subsequent release (or a manual step) is used to remove the deprecated schema elements.
Implementation Strategy
Define your migration as a Job with the appropriate annotations. Ensure the hook waits for completion so Helm does not start the rollout if the migration fails:
apiVersion: batch/v1
kind: Job
metadata:
name: db-migration
annotations:
helm.sh/hook: pre-upgrade,pre-install
helm.sh/hook-weight: "5"
helm.sh/hook-timeout: 5m
spec:
template:
spec:
containers:
- name: migration
image: my-migration-tool:latest
command: ["migrate", "up"]
restartPolicy: OnFailure
Assumptions and Uncertainty
This strategy assumes your database supports transactional DDL (like PostgreSQL) or that your migration tool is idempotent. If the database does not support atomic migrations, a failed hook execution could leave the schema in an inconsistent state, requiring manual intervention. We also assume your application code is designed to handle "missing or extra columns" gracefully during the transition period.
Verification
To verify this setup works, trigger a helm upgrade that adds a required column. Monitor the logs of the existing pods to ensure they do not throw errors when the migration job completes successfully.