C-order or F-order: which layout for column-wise NumPy reductions?
0 reputation · 27 Apr 2025, 01:56 UTC
C-order vs F-order for a column-heavy reduction workload
Consider a NumPy 2.x design where large 2-D float64 arrays, too big to stay in CPU cache, are produced as C-order row-major buffers. The dominant cost is a reduction along axis 0 (column-wise), repeated many times, with only occasional row-wise passes along axis 1.
The documented trade-off is clear at the layout level: C-contiguous memory favors row-wise traversal, F-contiguous memory favors column-wise traversal, and np.asfortranarray returns an O(1) view only when the input is already F-contiguous — otherwise it copies. At this scale, one copy may be negligible next to many repeated reductions, or it may dominate if the reduction runs only a few times. Internal blocking in NumPy's reductions may also narrow the stride penalty, making the real gap hard to predict from layout rules alone. The current layout of any array is directly observable through its C_CONTIGUOUS and F_CONTIGUOUS flags.
Specific questions:
- Does F-order contiguity reliably accelerate axis-0 reductions enough to justify the conversion copy for cache-exceeding arrays?
- How much do the occasional axis-1 passes and element-wise ufuncs slow down under F-order, and can that offset the gain?
- Is staying C-order and relying on NumPy's internal buffering a dependable middle path, or does it forfeit most of the column-wise benefit?