Diagnosing Slow Fortran Array Operations: Contiguity, Temporaries, and Vectorization
A step‑by‑step diagnostic guide for Fortran array performance: detect temporaries, enforce contiguous layout, enable vectorization, and verify with assembly and perf counters.
18 Sept 2025, 10:10 UTC

Recognizable Condition
Numerical kernels that manipulate large arrays run slower than expected, memory usage spikes during array slices, and the compiler does not emit SIMD instructions for loops that appear vectorizable.
Cause / Diagnostic Table
| Observed Symptom | Likely Cause | Quick Test |
|---|---|---|
| High allocation rate in profiler | Array slicing creates temporary copies | Compile with -fcheck=array_temps and watch for warnings |
No vector instructions in objdump -d | Missing CONTIGUOUS attribute or insufficient optimization flags | Recompile with -O3 -ftree-vectorize -march=native and inspect assembly |
| Cache‑miss heavy profile | Column‑major layout accessed row‑wise | Transpose loops or use reshape to match access pattern |
| Memory not released after procedure exit | Automatic deallocation disabled or allocatable without CONTIGUOUS | Enable -fcheck=all and verify deallocation points |
Ordered Checks
- Detect temporary arrays – Add
-fcheck=array_temps(GNU Fortran ≥10) to the compile line. Run the program; any “temporary array created” warning pinpoints slicing sites.
Run on a development box; no special permissions required. Warning output appears on stderr.gfortran -O2 -fcheck=array_temps -c kernel.f90 - Verify memory layout – Ensure every dummy argument and allocatable array that participates in heavy computation carries the
CONTIGUOUSattribute (Fortran 2008). Example:subroutine matmul(a, b, c) real, intent(in), contiguous :: a(:,:), b(:,:) real, intent(out), contiguous :: c(:,:) ... end subroutine - Enable aggressive vectorization – Recompile with
-O3 -ftree-vectorize -march=native. Then inspect the generated assembly for the hot loop:
Presence of AVX/AVX2/YMM registers confirms vectorization.gfortran -O3 -ftree-vectorize -march=native -S kernel.f90 -o kernel.s grep -n "vmovapd\|vaddpd\|vmulpd" kernel.s - Check column‑major access – Fortran stores arrays column‑major. If the innermost loop strides the first index, rewrite the loop nest or transpose the data once.
- Confirm automatic deallocation – Compile with
-fcheck=alland run under Valgrind or AddressSanitizer (-fsanitize=address) to catch leaks from allocatable arrays that lackCONTIGUOUS.
Fixes Tied to Findings
- Temporary arrays from slicing – Replace
a(:, i) = b(:, i) * 2.0with an explicit loop or ado concurrentblock that the compiler can fuse, eliminating the slice temporary. - Missing CONTIGUOUS – Add the attribute to dummy arguments and to allocatable declarations:
real, allocatable, contiguous :: work(:). This guarantees stride‑1 layout for the optimizer. - Insufficient compiler flags – Use the flag set from step 3 for production builds. For Intel ifx, replace with
-O3 -xHost -qopt-zmm-usage=high. - Row‑wise access on column‑major data – Either transpose the matrix once before the kernel (
call transpose(in, out)) or reorder loops so the inner index is the first dimension. - Memory not released – Ensure every allocatable array is either automatically scoped (function result, block construct) or explicitly deallocated. The
CONTIGUOUSattribute does not affect deallocation but helps the runtime track the descriptor.
Concrete Example: Matrix Multiply Before / After
Original kernel (slow, creates temporaries):
subroutine mmul_slow(a, b, c, n)
integer, intent(in) :: n
real, intent(in) :: a(n,n), b(n,n)
real, intent(out) :: c(n,n)
c = matmul(a, b) ! may allocate temporary for each slice
end subroutine
Optimized kernel (contiguous, vectorized):
subroutine mmul_fast(a, b, c, n)
integer, intent(in) :: n
real, intent(in), contiguous :: a(n,n), b(n,n)
real, intent(out), contiguous :: c(n,n)
integer :: i, j, k
do j = 1, n
do i = 1, n
c(i,j) = 0.0
do k = 1, n
c(i,j) = c(i,j) + a(i,k) * b(k,j)
end do
end do
end do
end subroutine
Compile both with gfortran -O3 -ftree-vectorize -march=native -S and compare the inner‑loop assembly. The fast version should show YMM registers and no calls to _gfortran_st_write for temporaries.
Limitations & Verification
- Fortran 77 code cannot use
CONTIGUOUSor allocatable arrays; refactoring may be required. - Vectorization success varies by compiler version (gfortran 12+, ifx 2023+, NVFortran 23+). Always verify with assembly.
-fcheck=array_tempsadds runtime overhead; disable for production builds.
Practical verification: Run the kernel under perf stat -e cycles,instructions,cache-references,cache-misses ./app before and after fixes. A ≥2× reduction in cycles per iteration with a corresponding rise in instructions‑per‑cycle confirms the diagnosis.
Escalation Criteria
- After applying all fixes, the hot loop still shows scalar instructions in assembly.
- Memory usage remains high despite
CONTIGUOUSand proper deallocation. - Performance gap >30 % compared to a vendor‑tuned BLAS call (e.g.,
dgemm).
If any criterion is met, profile with a hardware‑counter tool (VTune, perf record) and consider linking against an optimized BLAS/LAPACK library or rewriting the kernel in C/C++ with explicit intrinsics.
Diagram Labels
- Array slicing → temporary copy
- CONTIGUOUS attribute → aligned stride‑1 memory
- Compiler flags → vectorized loop body
- Performance measurement → speedup confirmation
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.