Diagnosing memcached slab-class fragmentation when hit ratio drops despite free bytes
Hit ratio falls, evictions rise, and free bytes still appear healthy? Slab-class fragmentation is usually the cause. This guide shows how to inspect per-class stats, identify the root, and apply targeted fixes.
08 Aug 2025, 06:55 UTC

Recognizable condition
\nYou check your monitoring dashboard and see that memcached’s hit ratio has dropped sharply, evictions are climbing, and request latency is spiking. Meanwhile, the bytes counter still shows room under the memory limit, and no out-of-memory crash appears. The pattern often follows a workload change, a restart, or a shift in key or value sizes. In short: the cache appears to have memory, but the slab allocator can’t put new objects where they’re needed, so it evicts under the least-recently-used (LRU) policy while other slab classes retain free chunks.
\nQuick takeaway
\nIf hit ratio falls, evictions rise, and bytes free looks healthy, suspect slab-class fragmentation before assuming overall memory exhaustion. Inspect per-slab utilization before adjusting capacity.
\nCause diagnostic map
\nMemcached groups items into slab classes by size. Each class has a fixed chunk size, and the LRU manager evicts from full classes. When key or value sizes change, items migrate to larger slabs, leaving other classes underutilized or overfull. Three common scenarios produce the observed symptoms:
\n| Cause | \nLikely symptom | \nFirst check | \n
|---|---|---|
| Hot slab class full, other classes idle | \nHigh eviction rate in one class, free chunks elsewhere | \nstats slabs shows non-zero evictions and low free_chunks in one class while others have the opposite | \n
| Item size growth moves keys into larger slabs | \nAverage object size per class rises, free chunks shrink in the new class | \nstats slabs reveals a class with larger chunk_size and fewer free chunks relative to item count | \n
| Restart creates cold slabs with uneven warm‑up | \nNew items fill a subset of classes; old classes retain free chunks from the warm‑up gap | \nuptime since restart correlates with class fill rates; recent stats shows gradual class occupation | \n
Ordered checks
\n- \n
- \n
Confirm overall health
\nUse your preferred method (telnet, nc, or a client library) to run:
\n
\necho 'stats'Record these values (replace CACHE_HOST and PORT with your deployment details):
\n- \n
- curr_items – current item count \n
- total_items – maximum items configured \n
- evictions – total evictions since start \n
- bytes and limit_maxbytes – used bytes versus the memory cap \n
- conn_cur – current connections \n
Where to run: on the memcached host or any machine that can reach the memcached text protocol port (default 11211). Required permissions: read‑only access; the stats command does not modify state. Expected checks: verify that bytes is well below limit_maxbytes while evictions rise, and that connections are within expected ranges.
\n \n - \n
Inspect per-slab utilization
\nRun:
\n
\necho 'stats slabs'Each line reports a slab class with:
\n- \n
- chunk_size – fixed size for items in this class \n
- total_chunks – total chunks allocated for this class \n
- used_chunks – chunks currently holding items \n
- free_chunks – empty chunks available \n
- items – number of items in the class \n
- evictions – items evicted from this class since start \n
Identify classes where evictions > 0 while free_chunks > 0 in other classes. This imbalance is the hallmark of slab‑class fragmentation.
\n \n - \n
Compare item count and average object size per class
\nDivide items by used_chunks to estimate average occupancy. If one class has many items but few free chunks, and another has many free chunks but few items, a size mismatch is likely. Compare the chunk_size values across the classes you identified.
\n \n - \n
Review recent changes
\nCorrelate the timeline with deployments, key‑size migrations, TTL changes, or auto‑scaling events that altered the value‑size distribution. A recent restart, a new deployment that started emitting larger payloads, or a change in cache‑key structure are common triggers.
\n \n
Fixes tied to findings
\n- \n
- \n
Normalize value sizes:
\nIf your workload shifted to larger objects, consider partitioning keys or adjusting application logic to keep values under the smallest slab chunk size (typically 96 bytes on 64‑bit builds). This reduces slab spread and keeps hot classes compact.
\n \n - \n
Increase memory or adjust slab-growth factor:
\nWhen the size distribution is broad and cannot be normalized, add memory so each class has enough chunks. Alternatively, start memcached with a smaller -f (slab growth factor) to produce more granular slab classes. Default is 1.25; lowering it creates more classes but may increase metadata overhead. Risk: adjusting memory or the growth factor requires a restart, which cold‑starts the cache and temporarily increases backend load. Schedule such changes during a maintenance window.
\n \n - \n
Enable slab rebalancing (where supported):
\nSome newer memcached builds support moving free chunks between classes without a full restart. Check your version’s man page for slab rebalance or proxy‑driven rebalancing capabilities. Caution: slab rebalancing behavior is version‑sensitive and may be unavailable or behave differently across releases.
\n \n - \n
Flush or warm specific slabs after a restart:
\nIf a recent restart left slabs cold, use a controlled warm‑up (e.g., pre‑load the most‑frequently accessed key subset) or a targeted flush_all followed by a gradual cache repopulation to avoid a thundering herd on backend services. Note: flush_all immediately invalidates every item in the cache; use only when you can absorb the backend load increase.
\n \n
Escalation criteria
\nElevate when any of the following hold:
\n- \n
- Sustained eviction rate exceeds your baseline while backend latency breaches the SLO, and per-slab inspection shows no improvement after rebalancing or size normalization. \n
- Repeated slab imbalance persists after attempted rebalancing or growth‑factor adjustment, indicating a fundamental mismatch between workload size distribution and memcached’s allocator parameters. \n
- Unexplained connection resets or item loss occur after a config change requiring restart during peak traffic, suggesting the restart window overlapped with traffic spikes and introduced transient state corruption. \n
In these cases, open a ticket with your platform team, include the full stats slabs output spanning the issue window, and note any recent config or deployment changes.
\nVerification and limitations
\nAfter applying a fix, repeat the ordered checks from the “Ordered checks” section. Compare the hit ratio, eviction count, and per-slab free‑chunk distribution against the pre‑fix baseline. A successful remediation will show evictions trending toward zero in previously hot classes and a more even free‑chunk spread across classes.
\n- \n
- Version sensitivity: Slab rebalancing and growth‑factor behavior differ between memcached releases. What works on 1.6.x may not apply—or may behave differently—on 1.5.x or binary‑protocol‑only builds. \n
- Protocol and proxy layers: The stats text output assumes a direct memcached connection. Behind a proxy (e.g., Envoy, HAProxy) or using the binary protocol, counters can be aggregated, renamed, or hidden. \n
- Restart cost: Adjusting memory or slab parameters requires a restart, which cold‑starts the cache and temporarily increases backend load. Schedule such changes during a maintenance window or alongside traffic ramp‑down. \n
Practical check: run echo 'stats slabs' before and after the change. Capture the output in a file, diff the class‑level free_chunks and evictions, and confirm the hit ratio stabilizes within your SLO window.
\n0 replies
A thoughtful contribution can make all the difference. Be the first to share one.