Recovery Procedure
There is no native rollback or downgrade command for DynamoDB Global Tables. Version upgrades are one-way operations and once a table has been partially upgraded, the version cannot be reverted in place. The supported recovery path is Point-in-Time Recovery to a timestamp before the upgrade attempt began, then rebuild Global Table replication on the restored resource.
Because a PITR restore creates a new table rather than modifying the existing one, a safe transition requires creating the restore, re-establishing replicas, shifting traffic, and then decommissioning the failed table.
- Confirm current state. Call DescribeTable in each region and check TableVersion and TableStatus. If any replica is UPDATING and stuck, management operations may be blocked and AWS Support intervention is required to force a state change.
- Verify PITR was enabled. PITR must have been enabled before the upgrade started. If enabled, choose a restore point immediately before the upgrade began. If not enabled, the only option is the most recent on-demand backup, with greater data loss.
- Create PITR restore. Restore the primary region table to a new table name using the pre-upgrade timestamp. The restore produces a new table ARN.
- Rebuild Global Table. The restored table is regional. Create a new Global Table using the restored table as the source and add the required replica regions to match the original architecture. Do not attempt to modify the failed Global Table in place.
- Shift traffic and update references. Update application configuration and any IAM policies, KMS grants, and CloudFormation stacks that reference the old table ARN to point to the new table. Use a blue/green cutover to limit downtime.
- Decommission. After verification, delete the failed Global Table to avoid duplicate billing.
Consistency and Data Verification
Restoring to a pre-upgrade point loses any writes made between the restore timestamp and the failure. Verification is therefore about confirming completeness of the restored baseline and identifying the loss window.
- Check Table version attribute via DescribeTable to confirm the new Global Table is on the intended version across all regions.
- Compare a sample of high-velocity partition keys between the failed table and the restored table to identify missing records in the loss window.
- Review CloudWatch metrics ReplicationLatency and PendingReplicationCount on the failed table to assess if data was stuck in transit during the partial upgrade.
- If DynamoDB Streams were enabled before the failure, the stream on the failed table can be inspected to understand writes that will not exist on the restored table.
Assumptions and Risks
This approach assumes PITR was enabled prior to the upgrade. A PITR restore changes the Table ARN, which breaks any resource-specific IAM policies or infrastructure as code references. Deleting the original table causes immediate downtime unless traffic is shifted first. Global Table replication may behave inconsistently if the upgrade failed during propagation, so manual rebuild is required.
One diagnostic detail that changes the recommendation: Is the table currently stuck in UPDATING state or ACTIVE but version-inconsistent across regions? A stuck UPDATING state may require AWS Support to release the table before PITR or replica operations can proceed.