Diagnosing and Fixing Azure Cosmos DB Hot Partitions: A Step-by-Step Guide
A diagnostic guide for Azure Cosmos DB hot partitions: recognize throttling on a single logical partition despite healthy account-level metrics, identify the skewed partition key pattern, run ordered verification queries, apply targeted key redesigns, and know when to escalate.
28 Jan 2026, 09:41 UTC

The Problem: Throttling on a Single Partition While Overall Throughput Looks Healthy
You provisioned 50,000 RU/s for your Cosmos DB container, but your monitoring shows only 15,000 RU/s consumed. Yet your application logs are filling with HTTP 429 (Too Many Requests) errors. The culprit is almost always a hot partition—a single logical partition receiving disproportionate traffic or storage, causing its dedicated physical partition to hit the 10,000 RU/s per-partition limit while the rest of your provisioned throughput sits idle.
Takeaway: Hot partitions manifest as localized throttling (429s) and latency spikes (p99 > 500 ms) tied to specific partition key values, even when account-level metrics look normal. The fix requires identifying the skewed key, redesigning it, and migrating data—there is no in-place ALTER PARTITION KEY.
Recognizable Symptoms
- Sudden 429 responses on a subset of requests while
Total Request Unitsmetric remains well below provisioned throughput. - Latency spikes (p99 > 500 ms) that correlate with specific partition key values in your logs.
- Storage skew > 10 GB on a single logical partition visible in Azure Monitor
PartitionKeyStatistics.
If you see these patterns, stop scaling throughput—it won't help. The bottleneck is physical partition saturation, not account capacity.
Cause and Diagnostic Quick-Reference
| Pattern | Typical Key Example | Why It Creates a Hot Spot |
|---|---|---|
| High cardinality but uneven access | tenantId where one tenant generates 80% of traffic | One logical partition absorbs the dominant tenant's RU/s |
| Low cardinality | status with values 'active'/'inactive' | All 'active' writes hit a single partition |
| Monotonically increasing key | timestamp or auto-incrementing ID | New writes always target the newest logical partition |
| Fan-out queries missing partition key filter | Query SELECT * FROM c WHERE c.type = 'order' without partition key | Cross-partition query overhead mimics partition skew; every partition pays RU cost |
Ordered Verification Steps
Run these checks in sequence. Each narrows the root cause before you commit to a migration.
1. Enable Diagnostic Logs and Query Partition-Level Metrics
Where: Azure Portal → Cosmos DB account → Diagnostic settings → Send to Log Analytics workspace.
Permissions: Contributor on the Cosmos DB account and Log Analytics workspace.
Query in Log Analytics:
AzureDiagnostics
| where ResourceProvider == "MICROSOFT.DOCUMENTDB" and Category == "PartitionKeyRUConsumption"
| summarize max(RUConsumption) by PartitionKeyValue, bin(TimeGenerated, 5m)
| where max_RUConsumption > 8000 // 80% of 10k RU/s per-partition limit
| order by max_RUConsumption desc
Expected check: Top partition key values consuming > 8,000 RU/s sustained over 5-minute windows.
Risk: Diagnostic log ingestion costs; filter to PartitionKeyRUConsumption category only to control spend.
2. Surface Top Partitions by Document Count and Size
Where: Run via Azure Cosmos DB Data Explorer or SDK (Python, .NET, Node.js) against the target container.
Permissions: DocumentDB Account Contributor or custom role with Microsoft.DocumentDB/databaseAccounts/readMetadata and data-plane read.
Query:
SELECT c._partitionKey, COUNT(1) AS docCount, SUM(c._ts) AS approxSize
FROM c
GROUP BY c._partitionKey
ORDER BY docCount DESC
LIMIT 10
Placeholders: None—_partitionKey and _ts are system properties.
Expected check: Top partition > 5% of total documents or > 2 GB storage signals skew.
Risk: This query scans the entire container; run during low-traffic windows or on a staging copy.
3. Correlate 429 Responses with Partition Key Values
Where: Application logs (App Insights, custom logging) capturing SDK response headers.
What to capture: RequestCharge, ActivityId, and the partition key value used in each request.
Diagnostic decision: If 429s cluster on specific key values matching step 1 or 2, you have confirmed the hot partition. If 429s are distributed but queries lack partition key filters, see cause #4 in the table above.
4. Verify Indexing Policy Includes Partition Key Path
Where: Azure Portal → Container → Scale & Settings → Indexing Policy, or SDK ContainerProperties.IndexingPolicy.
Check: Ensure includedPaths contains "/partitionKeyPath/*" (or your specific path) and no "/*" exclusion accidentally removes it. Missing index on the partition key forces scans that amplify RU consumption on the hot partition.
Targeted Fixes Mapped to Root Causes
Each fix requires a new container and data migration—Cosmos DB does not support changing a partition key in place.
Cause 1: Uneven Access on High-Cardinality Key (e.g., tenantId)
Fix: Introduce a synthetic suffix to spread the dominant tenant across multiple logical partitions.
// Before: partitionKey = "tenant_123"
// After: partitionKey = "tenant_123_" + hash(userId) % 100
// Example: "tenant_123_42"
Alternative: Hierarchical key tenantId/departmentId if departments have balanced workloads.
Constraint: Total key length < 2 KB. Hash suffix adds ~3–4 chars; well within limit.
Cause 2: Low-Cardinality Key (e.g., status)
Fix: Combine with a high-cardinality attribute.
// Before: partitionKey = "active"
// After: partitionKey = "active_" + correlationId
// Example: "active_550e8400-e29b-41d4-a716-446655440000"
Trade-off: Queries filtering only on status become cross-partition. Create a materialized view container with status as partition key if that query pattern is critical.
Cause 3: Monotonically Increasing Key (e.g., timestamp)
Fix: Prepend a hash prefix to distribute writes across partitions while preserving rough time ordering.
// Before: partitionKey = "2026-10-09T14:30:00Z"
// After: partitionKey = hash(timestamp) + "_" + timestamp
// Example: "a3f2_2026-10-09T14:30:00Z"
Migration path: Use change feed processor to backfill existing documents with new key format into new container. Allow 10–30 minutes after cutover for physical partition splits to complete before re-evaluating.
Cause 4: Fan-Out Queries Without Partition Key Filter
Fix: Rewrite queries to include the partition key, or create a materialized view container keyed for the query pattern.
// Before (cross-partition):
SELECT * FROM c WHERE c.type = 'order'
// After (single-partition):
SELECT * FROM c WHERE c.type = 'order' AND c.tenantId = 'tenant_123'
// Or create materialized view container with partitionKey = type
Verification: Compare RU charge before/after using SDK RequestCharge header. Expect 10–100x reduction for targeted queries.
Escalation Criteria: When to Open a Severity-A Support Ticket
Engage Azure Support immediately if any condition persists:
- Storage limit breach: Any single logical partition exceeds 20 GB (hard limit; physical partition cannot split further).
- Sustained throughput saturation: Single partition sustains > 10,000 RU/s for > 15 minutes after key redesign deployment (physical partition split cannot keep up).
- Replication impact: Cross-region replication lag > 5 minutes due to hot partition in write region.
Before escalating, confirm you have allowed 10–30 minutes for automatic physical partition splits after traffic redistribution.
Limitations and Practical Verification
- No in-place key change: Every fix requires new container + migration. Plan for dual-write or change-feed-based cutover.
- Hash suffixes sacrifice range queries: Original attribute (timestamp, tenantId) loses efficient range scans. Plan secondary indexes or materialized views for those patterns.
- Serverless accounts differ: 1 MB per logical partition limit; this guide applies to provisioned-throughput mode only.
- Diagnostic log costs: Filter to
PartitionKeyRUConsumptioncategory; disable after investigation.
Ongoing Validation Checklist
- Weekly: Run the GROUP BY partition key query; alert when top partition > 5% of documents or > 2 GB storage.
- Continuous: Azure Monitor alert on
PartitionKeyRUConsumption> 80% of 10k RU/s for 5 consecutive minutes. - Pre-deployment: Simulate production key distribution in staging container; confirm 429 rate drops to < 0.1%.
- Post-migration: Monitor change-feed processor lag; should return to < 1 s end-to-end.
- Billing: Compare RU/s consumed vs. provisioned before/after; expect 20–40% reduction in wasted RU from throttling retries.
Hot partitions are a design-time problem that surfaces at runtime. The diagnostic steps above let you confirm the root cause before committing to a migration, and the verification checklist prevents recurrence.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.