Implementing Terraform State Locking with DynamoDB
Prevent state corruption in AWS by implementing Terraform state locking with DynamoDB. Learn the minimal design, IAM requirements, and how to handle stale locks.
13 Jul 2025, 19:43 UTC

The Problem: Concurrent State Corruption
When multiple engineers or CI/CD pipelines execute terraform apply or terraform destroy simultaneously against the same state file, they risk creating a race condition. Without a locking mechanism, two processes may read the same state, calculate different changes, and attempt to write back conflicting updates. This leads to state corruption, orphaned resources, and inconsistent infrastructure.
The takeaway: To prevent concurrent modifications, you must implement a distributed locking mechanism. In AWS environments, the standard approach is using a DynamoDB table to act as a mutex (mutual exclusion) lock for the S3 backend.
Minimal Design Requirements
A state locking architecture requires a backend that supports locking and a separate database to track the lock status. For AWS, this is the S3 backend paired with a DynamoDB table.
The smallest suitable design consists of:
- S3 Bucket: Stores the
.tfstatefile. - DynamoDB Table: A single table used exclusively for locking.
- Primary Key: The table must have a partition key named
LockIDof typeString.
Terraform does not require any other columns to function, though adding a TTL (Time to Live) attribute can help automate the cleanup of stale locks if a process crashes before releasing the lock.
Implementation and Configuration
The DynamoDB table must be created before the backend can be initialized. This is typically done via a separate bootstrap process or the AWS Console.
Table Configuration
Run this command via the AWS CLI to create the required table. This example uses on-demand capacity to avoid throttling during bursts of CI/CD activity:
aws dynamodb create-table \
--table-name terraform-state-lock \
--attribute-definitions AttributeName=LockID,AttributeValueType=S \
--key-schema AttributeName=LockID,KeyType=HASH \
--billing-mode PAY_PER_REQUEST
Terraform Backend Block
Configure the backend in your Terraform files. This tells Terraform to check the DynamoDB table before any operation that modifies state.
terraform {
backend "s3" {
bucket = "my-terraform-state-bucket"
key = "prod/terraform.tfstate"
region = "us-east-1"
dynamodb_table = "terraform-state-lock"
encrypt = true
}
}
Trust and Data Boundaries
The lock table is a critical piece of infrastructure. If an unauthorized user can delete items from the table, they can bypass locking and corrupt your state.
IAM Boundary: The IAM role executing Terraform should be granted the minimum permissions required to manage the lock. Avoid dynamodb:*. Instead, use a scoped policy:
dynamodb:GetItem: To check if a lock exists.dynamodb:PutItem: To acquire a lock.dynamodb:DeleteItem: To release the lock after completion.
These permissions should be restricted specifically to the ARN of the terraform-state-lock table.
Operational Checks and Failure Modes
Locking is not a "set and forget" feature; it requires monitoring to ensure pipeline health.
Verification Process
To verify the lock is working, open two terminal windows. Run terraform apply in the first. While it is planning or applying, run terraform apply in the second. The second process should fail immediately with a Lock Failed error, stating that the state is locked by another process.
Failure Modes
| Scenario | Effect | Resolution |
|---|---|---|
| Process Crash | The lock remains in DynamoDB (stale lock). | Use terraform force-unlock [LOCK_ID] after verifying no other process is running. |
| DynamoDB Throttling | Terraform cannot acquire the lock and fails. | Switch to On-Demand capacity or increase Provisioned Throughput. |
| Network Partition | Timeout during lock acquisition. | Check connectivity to the DynamoDB endpoint; implement retry logic in CI/CD. |
Conditions for Redesign
The DynamoDB/S3 pattern is sufficient for most AWS-native workloads. However, you should reconsider this design if:
- Multi-Cloud Requirements: If you move to Azure or GCP, you must migrate to native locking (e.g., Azure Blob Storage leases).
- Managed State: If you adopt Terraform Cloud or Terraform Enterprise, locking is handled internally by the platform, and the DynamoDB table becomes redundant.
- Global Scale: If you require cross-region disaster recovery for the state lock itself, you would need to implement DynamoDB Global Tables, though this adds significant complexity to lock consistency.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.