Choosing Aerospike Storage Engine: Data-in-Memory vs Hybrid Memory
A decision guide for choosing between Aerospike Data-in-Memory and Hybrid Memory storage engines, covering persistence, latency, capacity, strong consistency, and operational constraints with concrete configuration examples and validation commands.
16 Sept 2025, 00:18 UTC

The Decision You Can’t Reverse
Aerospike namespaces lock in a single storage-engine at creation time: memory (Data-in-Memory) or device (Hybrid Memory). Changing it later requires a full namespace rebuild and data migration via dual-write or XDR. The choice dictates persistence guarantees, capacity limits, latency profiles, and whether Strong Consistency (SC) is even available. This guide walks through the constraints, compares the two engines, and shows how to validate your selection before you commit.
Quick Comparison
| Attribute | Data-in-Memory (memory) | Hybrid Memory (device) |
|---|---|---|
| Data location | RAM (indexes + records) | SSD (records), RAM (primary indexes ≈64 B/record) |
| Persistence | Optional snapshots; with persistence enabled writes to memory‑mapped file | Durable on every write when commit-to-device true (default) |
| Strong Consistency | Not supported | Supported (since 5.6) via Raft on master replica |
| Typical p99 latency | Sub‑millisecond, low variance | 0.2–0.5 ms added SSD read latency on NVMe |
| Capacity limit | RAM ≥ full dataset × replication factor | Dataset can exceed RAM; SSD capacity is primary limit |
| Defragmentation | None | Required; consumes SSD bandwidth and CPU |
| Migration cost | Full records move over network | Only indexes move; records stay on local SSD |
| Cloud block storage | Works on any volume | Requires local NVMe instance store or certified config; EBS/PD often violates raw‑device requirement |
Trade‑offs in Detail
Persistence and Durability
Data-in-Memory can survive a clean restart by loading a snapshot file, but any writes after the last snapshot are lost on crash unless you enable persistence (writes to a shared memory file). That file lives in the OS page cache, so memory pressure can cause latency outliers — monitor system_free_mem_pct to avoid hitting the eviction trigger. Hybrid Memory with commit-to-device true (the default) makes every write durable on the SSD immediately; no separate snapshot mechanism is needed.
Latency Profile
Data-in-Memory delivers consistent sub‑millisecond p99 latency because all reads and writes hit RAM. Hybrid Memory adds the SSD read path: on modern NVMe that’s typically 0.2–0.5 ms at p99. Write latency depends on commit-to-device; setting it to false acknowledges after writing to the write‑block buffer in RAM, but you lose durability on power loss. For latency‑sensitive workloads that can tolerate occasional data loss, that’s a knob; for most, keep it true.
Capacity Planning
Use these formulas to size each engine:
- Data-in-Memory RAM = (record_size + overhead) × records × replication_factor. Overhead includes index entry (~64 B) and internal metadata.
- Hybrid Memory RAM = (64 B × records × replication_factor) + index_overhead + sindex_overhead. Secondary indexes (sindex) live in RAM for both engines and can exceed memory limits even when data fits on SSD.
- Hybrid Memory SSD = (record_size + write_block_overhead) × records × replication_factor / defrag_threshold. The defrag threshold (default 50 %) determines how much free space must exist before defragmentation kicks in.
Strong Consistency
SC mode is only available with the device engine (Aerospike 5.6+). It provides linearizable reads and writes via Raft consensus on the master replica. The trade‑off: write latency increases because a quorum (e.g., 2 of 3 replicas) must acknowledge, and availability drops during network partitions. Ensure your replication factor and quorum size match your SLA — typically RF=3, quorum=2.
Operational Constraints
- Defragmentation (Hybrid Memory only) reclaims space from updates and deletes. It consumes SSD write bandwidth and CPU. Tune
defrag-lwm-pct,defrag-sleep, anddefrag-queue-minto prevent read latency spikes during heavy write workloads. Monitordefrag_qanddevice_write_bytes; keep SSD IOPS headroom at ≥3× sustained write throughput. - Migration/Rebalance Data-in-Memory moves full records across the network; Hybrid Memory moves only indexes, reducing network traffic but requiring SSD capacity headroom on each node to receive incoming partitions.
- Cloud Deployments Hybrid Memory’s raw‑device requirement often rules out network‑attached block storage (EBS, Persistent Disk). Use local NVMe instance store or Aerospike‑certified cloud configurations.
Concrete Configuration Examples
Data-in-Memory Namespace
namespace example {
memory-size 100G
storage-engine memory
data-in-memory true
persistence true
replication-factor 2
# optional: snapshot interval
# persistence-schedule 0 0 * * *
}
Run on a node with enough RAM for the entire dataset plus indexes. The persistence flag enables the memory‑mapped file; without it, a crash loses all data since the last snapshot.
Hybrid Memory Namespace with Strong Consistency
namespace example {
memory-size 32G
storage-engine device {
device /dev/nvme0n1
write-block-size 128K
}
strong-consistency true
replication-factor 3
commit-to-device true
defrag-lwm-pct 50
defrag-sleep 100
defrag-queue-min 10
}
The device stanza points to a raw NVMe partition (not a filesystem). write-block-size should match your SSD’s erase block (commonly 128 KB or 256 KB). Defrag parameters are starting points; adjust based on observed defrag_q.
Validation Steps
- Confirm storage engine on each node:
Run as root or a user withasadm -e "info namespace/example" | grep storage-engineasadmaccess. Output must showmemoryordevicematching your design. - Check index memory usage:
Verifyasadm -e "show index"index-used-bytes+sindex-used-bytes<memory-size×high-water-memory-pct(default 60 %). - For Hybrid Memory, validate SSD health and defrag lag:
Ensureasadm -e "show stat device" asadm -e "show stat defrag"defrag-qstays near zero under load anddevice_write_bytesdoesn’t saturate the drive’s rated IOPS. - Benchmark latency distribution with production‑like records:
Run from a client machine with network proximity to the cluster. Compare p99 against your SLA.asbench -h -n example -k 1000000 -w R,U -r 1000000 -o latency.txt - Simulate node failure (in a test cluster):
Observe migration progress. In SC mode, confirm no data loss; in AP mode, verify the divergence window is acceptable.kill -9 # on another node: asadm -e "show migrations"
Limitations and When to Re‑evaluate
- Changing storage engine after deployment is not an in‑place operation — plan for a full migration window.
- Hybrid Memory on non‑local block storage (EBS, PD) is unsupported and will cause unpredictable latency or data corruption.
- Secondary indexes can become the memory bottleneck even when data fits on SSD; monitor
sindex-used-bytesclosely. - Defragmentation stalls if SSD write bandwidth is exhausted; provision ≥3× sustained write throughput headroom.
Final Checklist
- Namespace
storage-enginematches design (memory vs device). - Index + sindex memory usage below high‑water mark.
- For device engine: SSD health OK, defrag queue near zero, IOPS headroom ≥3×.
- Benchmark p99 latency meets SLA under realistic load.
- Failure simulation shows expected migration behavior and no data loss (SC) or bounded divergence (AP).
If all checks pass, your storage engine choice is validated for production. Re‑run the benchmark and failure simulation after any major version upgrade or hardware change.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.