Speeding Up MATLAB Code with parfor: When and How to Use It
Learn how to decide if a loop is a good candidate for MATLAB's parfor, set up a parallel pool, and measure the actual speed‑up on a multicore laptop.
03 Mar 2026, 09:08 UTC

Problem: Your MATLAB script spends too long in a simple loop
You have a script that processes each element of a large vector independently—for example, computing a costly function for every point in a simulation grid. The straightforward for loop runs on a single core and leaves idle CPU cycles on your laptop or workstation.
Thesis: Use parfor when each iteration is independent and heavy enough to outweigh parallel overhead
MATLAB’s parfor (parallel for‑loop) distributes loop iterations across the workers in a parallel pool managed by the Parallel Computing Toolbox. If the work per iteration is substantial and there are no data dependencies, you can often see 2×‑8× speed‑up on a quad‑core machine.
Section 1: Checking readiness
Before rewriting the loop, verify three things:
- Toolbox license – run
license('test', 'Parallel_Toolbox'). It should return1. - Available workers – start a local pool and inspect its size:
% Run in the MATLAB command window or a script
pool = gcp('nocreate'); % get current pool if any
if isempty(pool)
pool = parpool('local'); % start a new local pool
end
fprintf('Pool has %d workers\n', pool.NumWorkers);
The ideal worker count equals the number of physical cores (feature('numCores')). If you see fewer workers, close other MATLAB sessions or adjust the pool size with parpool('local', N).
Section 2: Structuring the loop for parfor
parfor imposes a few constraints:
- Loop variable must be consecutive integers.
- Variables used inside the loop are classified as broadcast (read‑only), sliced (each worker gets a distinct portion), or reduction (combined across workers).
- You cannot grow an array or concatenate results inside the loop without preallocation.
Consider the example of computing the sum of squares of a large vector:
N = 10^7;
x = randn(N,1);
% Serial version
tic;
s = 0;
for i = 1:N
s = s + x(i)^2;
end
t_serial = toc;
fprintf('Serial time: %.3f s\n', t_serial);
To parallelise, we turn the accumulation into a reduction variable:
tic;
s_par = 0;
parfor i = 1:N
s_par = s_par + x(i)^2; % reduction variable
end
t_par = toc;
fprintf('parfor time: %.3f s\n', t_par);
fprintf('Speed‑up: %.2fx\n', t_serial/t_par);
MATLAB automatically slices x as a broadcast variable (read‑only) and treats s_par as a reduction, summing the partial results from each worker.
Section 3: Measuring and validating the gain
Use tic/toc around both versions, ensuring the pool is already running before the timed section to avoid counting pool startup overhead. A quick sanity check:
% Verify correctness
if abs(s - s_par) < 1e-10
fprintf('Results match within tolerance.\n');
else
error('Results differ – check for hidden dependencies.');
end
If the speed‑up is disappointing (<1.2×), consider:
- Increasing the work per iteration (e.g., more expensive function calls).
- Reducing data transfer: avoid large broadcast variables that are not needed.
- Checking for hidden dependencies such as file I/O or shared resources inside the loop.
Trade‑off and limitation
The main limitation is that parfor cannot accommodate loops where each iteration depends on the previous one or where you need to dynamically grow an output array. In those cases you must redesign the algorithm (e.g., use parfeval for asynchronous tasks or rewrite to use built‑in vectorised functions). Additionally, debugging is harder because breakpoints inside the worker code are not honored; you must rely on lasterror or afterEach for logging.
Actionable closing
Next time you spot a long‑running, embarrassingly parallel for loop, follow this checklist:
- Confirm the Parallel Computing Toolbox is licensed.
- Start a pool with
parpool('local')and note the worker count. - Ensure each iteration is independent and contains enough work to outweigh communication overhead.
- Refactor the loop to respect parfor’s variable classifications (pre‑allocate outputs, use reductions).
- Time the parallel version against the serial baseline and verify numerical equivalence.
- If the speed‑up meets your needs, keep the parfor; otherwise, reconsider the algorithm or look into GPU alternatives.
By treating the loop as a data‑parallel problem and letting MATLAB handle the worker management, you can often reclaim idle CPU cycles with only a few lines of code changes.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.