np.memmap vs np.load mmap_mode for writable large arrays under memory pressure
0 reputation · 21 Sept 2020, 08:57 UTC
Context
Working with arrays that exceed available RAM requires memory-mapped file access. NumPy provides two documented entry points: np.memmap for explicit control over offset, shape, and mode, and np.load with mmap_mode for a simpler loading API. Both avoid eager allocation for read-only workloads, but the constraint shifts when the array must be modified in place.
Constraint
The target environment has limited physical memory and a read-write dataset larger than RAM. The application needs deterministic persistence semantics: changes should reach durable storage without manual flush() calls scattered through the codebase, yet the process cannot afford a full in-memory copy.
Unresolved decision
NumPy does not standardize copy-on-write or automatic flushing guarantees for writable mappings. np.memmap exposes flush() but leaves flush timing and dirty-page propagation to the OS. np.load with mmap_mode='r+' or 'c' (copy-on-write) offers fewer parameters and its write-back behavior varies by platform and NumPy version. Neither approach documents a portable contract for when modifications become visible to other processes or survive a crash.
Which approach gives more predictable write persistence for a memory-constrained writer: the explicit np.memmap with manual flush discipline, or np.load with mmap_mode='c' relying on OS copy-on-write? Does either provide a documented guarantee that surviving pages reflect the last written state after an unclean shutdown?