Diagnosing and Fixing Memory‑Optimized Bucket Quota Exhaustion in Couchbase 7.x
When Couchbase clients see write latency spikes and time‑outs, the culprit is often a memory‑optimized bucket hitting its quota. This guide walks through the symptoms, diagnostic checks, and remediation steps to restore steady performance.
31 Jan 2026, 08:55 UTC

Recognizable Condition
Clients report intermittent write latency spikes, occasional timeout or out_of_memory errors, and a sudden rise in operation timeouts in application logs. Cluster‑wide metrics show the active eviction rate on data nodes rising above 0 % and the memory used approaching the configured memory quota for the bucket. CPU on the data service climbs during eviction periods.
Cause Diagnostic Table
| Potential Cause | What to Look For |
|---|---|
| Memory quota too small for working set | Memory used ≈ quota; active eviction > 0 % for > 5 min |
| Large value items or growing documents | Average item size > 200 KB; size distribution skewed to > 1 MB |
| Write‑heavy workload without TTL | Item count steadily increasing; no expiration set |
| Replica count higher than needed | Replica count per bucket > 1; each replica consumes full memory copy |
| Compaction or index pressure | Index memory usage high; data service memory drops after compaction |
Ordered Checks
- Verify per‑node memory usage. Run on a data node:
Check that# cbstats -u admin -p password -c node_ip | grep "data memory used" # cbstats -u admin -p password -c node_ip | grep "data memory quota"data memory usedis < 90 % ofdata memory quotaunder steady load. - Inspect eviction metrics. On the cluster console or via REST:
A non‑zerocurl -s -u admin:password http://node_ip:8091/pools/default/buckets/bucket_name | jq .evictionPolicy curl -s -u admin:password http://node_ip:8091/pools/default/buckets/bucket_name | jq .evictionRateevictionRateindicates active eviction. - Analyze document size distribution. Use the SDK or N1QL:
SELECT AVG(JSON_LENGTH(value)) AS avg_size, COUNT(*) AS count FROM `bucket_name` GROUP BY bucket_name LIMIT 1; - Check write throughput and TTL usage. Run a simple write‑rate test or review application logs for
setoperations withoutexpiry. - Confirm replica configuration. In the bucket settings or via REST:
curl -s -u admin:password http://node_ip:8091/pools/default/buckets/bucket_name | jq .replicaNumber - Review index and compaction impact. Monitor
index memory usageandcompactionevents in cluster logs.
Fixes Tied to Findings
- Increase memory quota or add data nodes. Adjust
memory quotain bucket settings or scale the cluster, then runrebalance. - Reduce large document sizes. Refactor payloads, enable compression, or move large blobs to a separate bucket or external storage.
- Introduce TTL or cleanup old data. Add
expiryon new writes or run a one‑time cleanup script to delete stale items. - Lower replica count. If durability requirements allow, reduce
replicaNumberto 1 or 0, then rebalance. - Adjust eviction policy. For memory‑optimized buckets,
value_onlyis the default; considerfull_evictionif the bucket holds very large values that can be evicted.
Escalation Criteria
- Eviction rate remains > 5 % after increasing quota and reducing data volume for > 15 min.
- Client errors persist with
disk_fullorout_of_memoryafter all memory‑level fixes. - Rebalance fails or nodes become unhealthy during quota changes.
- Durability constraints cannot be met due to reduced replicas, indicating a deeper architectural issue.
Practical Example
In a 3‑node cluster, bucket app_data is memory‑optimized with a 512 MB quota. After a spike in writes, the admin sees data memory used 490 MB and eviction rate 12 %. The following steps were taken:
- Run
cbstatson each node; confirm all nodes at 95 % usage. - Query average document size:
SELECT AVG(JSON_LENGTH(value)) AS avg_size FROM `app_data` LIMIT 1;returns 350 KB. - Reduce average size by splitting large JSON objects into two smaller documents and enabling compression.
- Add
expiryof 30 days to new writes. - Increase quota to 640 MB via the web console and trigger
rebalance. - After 30 min,
data memory useddrops to 420 MB andeviction rateis 0 %. - Monitor client latency: 95th percentile drops from 600 ms to < 50 ms.
Limitations
- Metrics names and units can vary between 7.x releases; verify against your cluster’s
poolsAPI. - Rebalancing can temporarily degrade performance; schedule during maintenance windows.
- Disk‑optimized buckets behave differently; this guide applies to memory‑optimized buckets only.
Verification Checklist
- Cluster reports
data memory used< 90 % of quota under steady load. - Active eviction rate is 0 % in the last 10 minutes.
- Client write latency percentiles return to baseline (< 50 ms for 95th percentile).
- Bucket item count and average size trend stabilizes or decreases.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.