Guide
Diagnosing DynamoDB Throttling and Hot Partition Issues
A diagnostic guide recognizing DynamoDB throttling and hot partition symptoms, ordered checks, capacity and key redesign fixes, and AWS Support escalation criteria.
Published by Tasadduq Burney
11 Jun 2026, 01:23 UTC
3 min57.8K views0

Recognizable condition
When DynamoDB returns ProvisionedThroughputExceededException or HTTP 400 errors, and CloudWatch shows a spike in ThrottledRequests, the table is experiencing either provisioned capacity limits or a hot partition. ProvisionedThroughputExceededException means the request rate exceeds the provisioned read or write capacity units. A hot partition occurs when traffic concentrates on a single partition key, causing disproportionate latency and error rates for that key.
Cause / Diagnostic table
| Symptom | Likely cause |
|---|---|
Broad increase in ThrottledRequests across all keys | Insufficient provisioned read/write capacity |
| Latency spikes for a specific partition key | Hot partition due to skewed access pattern |
| Uneven storage size shown in the DynamoDB console | Hot partition caused by non-uniform key distribution |
Ordered checks
- Open the CloudWatch console for the table and locate the metrics
ThrottledRequests,ConsumedReadCapacityUnits,ConsumedWriteCapacityUnits,ProvisionedReadCapacityUnits,ProvisionedWriteCapacityUnits. Compare the throttled request count to the provisioned capacity; ifThrottledRequests> 0 while consumed units are near provisioned units, capacity is the limiter. - Use DynamoDB Insights or the CLI to export item collection metrics and identify partition-key distribution:
aws dynamodb describe-table --table-name MyTable --query Table.ItemCollectionMetrics. This returns per-key request rates. Identify any partition key whose rate exceeds the average by a large factor (e.g., >5x). - Check the table storage size per partition key in the DynamoDB console under the Items tab to confirm skew.
Fixes tied to findings
- Insufficient capacity: Increase provisioned throughput. Example CLI (requires IAM permission
dynamodb:UpdateTable):
Run this in a terminal with AWS credentials that have the required permission. Expect a brief table availability interruption; test the change in a staging environment first. After running, monitoraws dynamodb update-table --table-name MyTable --provisioned-throughput ReadCapacityUnits=2000,WriteCapacityUnits=2000ThrottledRequestsin CloudWatch for a decreasing trend. - Enable auto scaling: Configure Application Auto Scaling with a target utilization (e.g., 70%). Required permissions:
application-autoscaling:RegisterScalableTargetandapplication-autoscaling:PutScalingPolicy. Example scalable-target registration:
Create a target-tracking policy that sets the desired capacity to maintain 70% utilization. Risk: scaling adjustments may lag behind sudden traffic bursts; combine with on-demand mode for spikes.aws application-autoscaling register-scalable-target --service-namespace dynamodb --resource-id table/MyTable --scalable-dimension dynamodb:table:ReadCapacityUnits --min-capacity 500 --max-capacity 5000 - Hot partition redesign: If a specific key (e.g., user_id) drives most traffic, add a random suffix or use a composite key. This requires migrating data to a new table with the new key schema. Define a new partition key named pk (string) and sort key named sk (string). Then copy items, updating the partition key to a value like pk = original_user_id-random_number(0,99) using AWS DMS or a custom script. Risk: table migration introduces downtime or data loss if not tested in staging first.
- On-demand mode: For highly unpredictable workloads, switch billing to PAY_PER_REQUEST:
Requires permission dynamodb:UpdateTable. Cost model changes; monitor daily spend after switching.aws dynamodb update-table --table-name MyTable --billing-mode PAY_PER_REQUEST
Escalation criteria
- Throttling persists (>0 ThrottledRequests) after at least one auto scaling adjustment cycle (≈5 minutes) despite capacity increases.
- Latency (measured via CloudWatch SuccessfulRequestLatency) exceeds 200 ms for more than 5 % of requests over a 10-minute window.
- Errors impact an SLA and remain unresolved after trying the above fixes.
- In these cases, open an AWS Support case, providing CloudWatch metric screenshots and the tables partition-key distribution report.
Limitations and verification
- Changing the partition key requires table migration; test the new key design in a staging environment to avoid data loss or downtime.
- Auto scaling may lag behind sudden bursts; combine with burst capacity or consider on-demand mode for traffic spikes.
- After applying any fix, monitor ThrottledRequests and latency for at least one full scaling cycle (typically 5 minutes) to confirm improvement.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.