MATLAB Vectorize vs Loop: What Modern Releases Actually Changed
The "always vectorize" advice is outdated. Modern MATLAB's JIT and implicit expansion make preallocated loops competitive. Profile first, then choose based on memory, readability, and your actual hardware.
16 Aug 2026, 19:22 UTC

The Problem: Outdated Advice Still Circulates
If you've written MATLAB for more than a few years, you've heard it: "Always vectorize. Loops are slow." That advice made sense in 2005. It's misleading today. The execution engine overhaul around R2015b and implicit expansion in R2016b changed the performance landscape, but many style guides and code reviews haven't caught up.
The engineering decision isn't vectorize versus loop. It's profile first, then choose the style that matches the problem and the hardware you're running on.
What Actually Changed
Two milestones matter. The R2015b execution engine rewrite introduced a just-in-time (JIT) compiler that optimizes plain indexed loops at runtime. Simple for-loops over preallocated arrays now run at speeds comparable to vectorized equivalents for many workloads.
R2016b added implicit expansion. Element-wise operators automatically broadcast compatible sizes, so Z = (X - mu) ./ sigma works when X is m-by-n and mu, sigma are 1-by-n row vectors. Before R2016b you needed bsxfun or repmat, which is why older codebases look different.
Preallocation remains the single most reliable loop optimization. Growing an array with x(end+1) = ... forces repeated reallocation and copying. zeros(n,m), cell(n,1), or string preallocation before the loop is still the first fix for slow loops, independent of the vectorize debate.
Worked Example: Column Normalization
Standardize each column of an m-by-n feature matrix X using per-column mean mu and standard deviation sigma. Both versions below assume mu and sigma are 1-by-n row vectors.
Preallocated Loop
function Z = normalize_loop(X, mu, sigma)
[m, n] = size(X);
Z = zeros(m, n); % preallocate
for k = 1:n
Z(:, k) = (X(:, k) - mu(k)) ./ sigma(k);
end
end
Vectorized with Implicit Expansion
function Z = normalize_vec(X, mu, sigma)
Z = (X - mu) ./ sigma; % broadcasts mu, sigma across rows
end
Both are a few lines. The vectorized version reads like the math. The loop version makes the column-wise operation explicit and writes directly into the preallocated output.
Memory and Readability Trade-offs
The vectorized expression allocates temporaries: X - mu creates an m-by-n array, then the division creates another. Peak memory can approach a few times the matrix size. The preallocated loop writes in place, so peak memory stays near one copy of Z.
On huge arrays the memory cost can cancel the speed win; on tiny arrays neither matters. MATLAB's column-major layout also means traversing down columns (first index varying fastest) is cache-friendly. The loop above does exactly that. A per-row loop or row-wise vectorization of the same math can be markedly slower.
Implicit expansion can silently broadcast mismatched shapes. Subtracting a column vector from a row vector yields an m-by-n matrix—a classic silent-bug source. Verify dimensions with size() in production code:
> size(rand(4,3) - rand(1,3))
ans =
4 3 % R2016b+ broadcasts row across columns
> size(rand(4,3) - rand(4,1))
ans =
4 3 % broadcasts column across rows
Profile-First Workflow
Don't optimize blindly. Use timeit for microbenchmarks—it handles warm-up and repeats better than raw tic/toc. Use profile on and the Profile Viewer to find real hotspots in your full pipeline before rewriting a normalization step.
% Microbenchmark in your release, on your hardware
f_loop = @() normalize_loop(X, mu, sigma);
f_vec = @() normalize_vec(X, mu, sigma);
t_loop = timeit(f_loop);
t_vec = timeit(f_vec);
fprintf('Loop: %.4f s, Vectorized: %.4f s\n', t_loop, t_vec);
Run this in your target MATLAB release. Results vary by release, OS, and hardware. Publish your own measured numbers with the release and hardware noted.
Where Each Style Wins
- Vectorization shines for element-wise math, reductions along a dimension, logical masking, and replacing tiny-loop function-call overhead.
- Loops win for sequential dependencies (recurrences, filters), I/O, incremental construction, and often readability when the algorithm is naturally iterative.
Actionable Closing
Next time you're tempted to vectorize a loop reflexively, stop. Run profile viewer on the full workflow. If the loop isn't a hotspot, leave it—readability wins. If it is, benchmark both styles with timeit on your release and hardware. Check memory with whos or the Memory Profiler if arrays are large. Document the choice with a comment linking to the benchmark results. That's the engineering decision, not a style rule from 2005.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.