Tuning Aerospike Hybrid Memory Architecture for Latency vs. Cost
Learn how adjusting the memory‑data ratio in Aerospike’s HMA affects read latency and SSD wear, with a step‑by‑step example and practical trade‑offs.
11 Jul 2026, 00:12 UTC

Problem: Balancing fast reads with affordable storage
Many Aerospike deployments keep the primary index in RAM while storing actual records either in memory or on SSD. If too little data resides in RAM, read latency rises because the server must fetch from flash. If too much data is kept in RAM, infrastructure costs climb and frequent migrations between tiers can increase SSD wear. The goal is to find a memory‑to‑data ratio that meets your read‑latency SLA without exceeding budget or causing excessive data movement.
Thesis: Adjust the memory‑data ratio via HMA to hit a latency target
Aerospike’s Hybrid Memory Architecture (HMA) lets you size the index and data tiers independently. By changing the namespace parameter that controls how much of the total data set lives in memory (often called memory-size or data-size depending on version), you directly influence the fraction of reads served from RAM versus SSD. Monitoring latency histograms and migration counters lets you iterate toward the optimal ratio.
Understanding HMA and the memory‑data ratio
The primary index (key → location pointer) always lives in RAM. Record bins can be placed in:
- Memory‑only: entire record in RAM (fastest reads, highest RAM cost).
- Hybrid: index in RAM, record data on SSD (lower RAM cost, higher latency for misses).
- Memory‑mapped SSD: Aerospike can also keep a portion of records in RAM while the rest spill to flash based on the configured ratio.
The memory‑data ratio is the proportion of the total data size that Aerospike attempts to keep in memory. For example, a ratio of 0.30 means Aerospike will try to keep roughly 30 % of the data set in RAM; the rest resides on SSD. This ratio is independent of the index size, which is always fully resident.
Worked example: shifting from 30 % to 60 % memory
Assume a namespace test with a total data size of 100 GB. The current configuration keeps 30 % (≈30 GB) in memory and 70 % on SSD. You observe a 99th‑percentile read latency of about 0.5 ms under a steady 10 k reads/sec workload with an 80 % hit‑rate.
To test whether increasing the memory fraction improves latency, follow these steps:
- Enable latency histograms on each node (requires admin privileges):
This starts collecting latency buckets without affecting traffic.asinfo -v 'latency:enable=true' - Record baseline metrics (run a read‑only benchmark with a known key set, e.g., using
aerospike-benchmark):
After the run, retrieve latency stats:aerospike-benchmark -h $HOST -n test -t 10 -R 10000 -l 1000000
Note the average and 99th‑percentile values.asinfo -v 'latency' - Check migration and defrag counters to ensure the current ratio isn’t causing excess movement:
Values under ~5 per second are generally healthy.asinfo -v 'statistics' | grep -E 'migrate|defrag' - Adjust the memory‑data ratio. In Aerospike 6.x the parameter is
memory-size(in GB). Edit the namespace stanza in/etc/aerospike/aerospike.conf:
This tells Aerospike to aim for ~60 % of the 100 GB data set in RAM.namespace test { memory-size 60G # changed from 30G storage-engine device {...} } - Perform a rolling restart** (one node at a time) to apply the change without downtime:
On each node, verify it rejoins the cluster before moving to the next.systemctl restart aerospike - Re‑run the same benchmark** and collect latency histograms again.
- Compare results**. If the 99th‑percentile latency drops toward the target (e.g., ~0.2 ms) and migration/defrag counters stay low, the new ratio is working. If counters spike, you may be causing too much data movement; consider a smaller increment or investigate write load.
Note: The latency numbers above are illustrative; actual outcomes depend on record size, write intensity, and SSD characteristics.
Trade‑offs and limitations
- Operational cost: Doubling the memory fraction roughly doubles RAM expenditure for the namespace.
- SSD wear: A higher memory ratio can increase write‑amplification if the workload is write‑heavy, as more records may migrate between RAM and SSD during defragmentation.
- Rolling restart required: Changing
memory-sizenecessitates a node restart; plan for a maintenance window or use a rolling upgrade strategy. - Granularity: The ratio is namespace‑wide; you cannot pin individual bins to memory or SSD without using separate namespaces or user‑defined functions.
Actionable closing
Start by measuring your current latency histogram and migration counters. Make small, incremental adjustments to the memory‑data ratio, restart nodes one at a time, and re‑measure. Stop when the 99th‑percentile read latency meets your SLA and migration/defrag counters remain low (< 5 /sec). Document the final ratio and keep the latency histogram enabled for ongoing visibility—this lets you catch drift caused by changing traffic patterns or SSD aging.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.