Diagnosing and Resolving Terraform State Lock Stalls with DynamoDB Backend
A step-by-step diagnostic guide for Terraform state lock stalls when using an S3 backend with DynamoDB locking. Covers how to identify stale locks via AWS CLI, verify no active runs, safely force-unlock, and prevent recurrence with lock_timeout.
19 May 2026, 02:55 UTC

Recognizable Condition
Terraform commands such as init, plan, or apply hang indefinitely or return Error acquiring the state lock. The error message typically includes a LockID that matches your workspace (for example, my-project/production) and a SessionID that identifies the process that originally acquired the lock. In a team environment, multiple engineers or CI pipelines report the same failure, blocking all infrastructure changes.
Cause and Diagnostic Quick Reference
| Observed Symptom | Likely Cause | Verification Step |
|---|---|---|
| Command hangs without timeout | DynamoDB lock item exists with no TTL; default backend timeout is 0 (infinite) | Query the lock table and inspect the Timestamp attribute |
Error acquiring the state lock with a SessionID from a terminated runner |
Previous Terraform process crashed or was killed before releasing the lock | Cross-reference SessionID with active EC2 instances, CI job logs, or local processes |
Lock Timestamp older than 30 minutes |
Stale lock left by a failed or interrupted run | Compare Timestamp to current time; if delta > safety threshold, treat as stale |
Multiple concurrent runs fail with same LockID |
Legitimate contention (another engineer is running apply) |
Confirm with team chat or CI dashboard before taking any unlock action |
Ordered Diagnostic Checks
- Identify the lock table name. It is defined in your backend configuration, e.g.,
dynamodb_table = "terraform-locks"in thebackend "s3"block. - Retrieve the lock item. Run the AWS CLI command below from a workstation with permissions to read the DynamoDB table (typically
dynamodb:GetItemon the lock table ARN). Replace placeholders with your values:
The response contains attributesaws dynamodb get-item \ --table-name terraform-locks \ --key '{"LockID":{"S":"my-project/production"}}'LockID,SessionID,Timestamp(ISO 8601 string), and optionallyOperationandInfo. - Check the lock age. Convert the
Timestampto epoch seconds and compare withdate +%s. A difference greater than 1800 seconds (30 minutes) is a strong indicator of a stale lock, assuming no long-running apply is expected.LOCK_TS=$(aws dynamodb get-item --table-name terraform-locks \ --key '{"LockID":{"S":"my-project/production"}}' \ --query 'Item.Timestamp.S' --output text) NOW=$(date -u +%s) LOCK_EPOCH=$(date -u -d "$LOCK_TS" +%s 2>/dev/null || date -u -j -f "%Y-%m-%dT%H:%M:%SZ" "$LOCK_TS" +%s) AGE=$((NOW - LOCK_EPOCH)) echo "Lock age: $AGE seconds" - Verify no active Terraform process holds the lock. The
SessionIDoften contains the hostname, PID, and a random suffix (e.g.,ip-10-0-1-42-1234-abc123). Search your CI system, EC2 console, or localps aux | grep terraformfor that identifier. If the originating process is confirmed terminated, proceed.
Fixes Tied to Findings
Stale Lock Confirmed (Age > Threshold, No Active Process)
Run terraform force-unlock with the exact LockID shown in the error or retrieved from DynamoDB. This command writes a ForceUnlock request to the backend, which removes the lock item. Execute from the same working directory that contains the backend configuration so Terraform resolves the correct lock table.
terraform force-unlock -force my-project/production
The -force flag skips the interactive confirmation. After the command succeeds, re-run the original Terraform command (plan or apply).
Legitimate Contention Detected
If another engineer or CI job is actively running, wait for it to complete. Do not force-unlock. Coordinate via your team's communication channel. If the run appears stuck (e.g., a provider timeout), the owner should cancel their Terraform process cleanly (Ctrl+C), which releases the lock automatically.
Recurring Stale Locks — Adjust Backend Timeout
Terraform's S3 backend with DynamoDB locking supports a lock_timeout argument (string, e.g., "20m"). When set, Terraform will abort lock acquisition after the duration instead of waiting indefinitely. Update your backend block:
terraform {
backend "s3" {
bucket = "my-terraform-state"
key = "my-project/production/terraform.tfstate"
region = "us-east-1"
dynamodb_table = "terraform-locks"
lock_timeout = "20m"
}
}
After changing the backend configuration, run terraform init -migrate-state to re-initialize. This does not modify existing state; it only updates the local backend metadata.
Escalation Criteria
- Force-unlock fails with a permission error: verify the IAM role/user has
dynamodb:DeleteItemon the lock table. Escalate to the AWS administrator. - Lock reappears immediately after force-unlock: a background process (e.g., a stuck CI agent) may be re-acquiring it. Hunt down the process by
SessionIDpattern. - State corruption suspected (plan shows unexpected diffs after force-unlock): stop, snapshot the S3 state object (
aws s3 cp s3://bucket/key key.backup), and engage a Terraform specialist before further writes. - DynamoDB table throttling (high
ConsumedReadCapacityUnitsorThrottledRequestsmetrics): consider provisioned capacity or on-demand mode; this is an infrastructure scaling issue, not a lock hygiene issue.
Verification After Resolution
- Run
terraform plan— it should complete without lock errors. - Confirm the lock item is absent or updated with a fresh
Timestampand your currentSessionIDby re-running theaws dynamodb get-itemcommand. - If you added
lock_timeout, trigger a concurrent test run (two terminals, same workspace) to verify the second run exits with a clear timeout message after ~20 minutes instead of hanging.
Limitations and Safety Notes
terraform force-unlockis a destructive operation. If two processes genuinely hold the lock (rare, but possible with clock skew or DynamoDB eventual consistency), forcing one out can leave the state in an inconsistent state. Always verify no active Terraform process before using it.- The
Timestampattribute is written by the locking client; a malicious or buggy client could write a future timestamp, making the lock appear fresh. Treat the timestamp as a heuristic, not a guarantee. - Lock timeout (
lock_timeout) only affects the acquisition wait. It does not impose a TTL on the lock item itself. DynamoDB TTL on the lock table is not supported by the Terraform backend. - This guide assumes the standard S3 + DynamoDB backend. Other remote backends (Consul, etcd, Azure Storage, GCS) have different lock semantics and CLI flags.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.