numpy.array() vs numpy.fromiter(): memory‑efficient streaming or fast bulk copy for large iterables?
0 reputation · 05 May 2023, 02:11 UTC
Goal
Convert a large Python iterable (e.g., a list with >10⁷ elements) to a NumPy ndarray while keeping peak memory usage low and avoiding unnecessary latency.
Constraints and uncertainty
numpy.array() always copies the input, doubling memory temporarily. numpy.asarray() avoids a copy only when the source is already an ndarray, which does not help for plain lists. numpy.fromiter() streams elements without storing the full input, but it requires the total count upfront for optimal allocation and may be slower per element. Version‑dependent optimizations in NumPy 1.26+ can reduce copy overhead, yet the documented default remains copy‑on‑create.
Open questions
- Does
numpy.fromiter()with a known length consistently use less peak memory thannumpy.array()for lists larger than 10⁷ elements? - How does the runtime of
fromiter()compare to the bulk copy path on NumPy 1.26+? - Are there cases where pre‑allocating an ndarray and filling it via
asarray()or direct assignment outperforms both approaches?