Accelerating Monte Carlo Simulations in MATLAB with parfor
Learn how to use MATLAB's parfor loops to accelerate Monte Carlo simulations, reducing computation time by distributing independent trials across multiple CPU cores.
03 Aug 2025, 07:21 UTC

The Bottleneck of Independent Trials
In engineering and financial modeling, Monte Carlo simulations are essential for estimating outcomes when variables are stochastic. Whether you are pricing a complex derivative or simulating structural fatigue, the process usually involves running millions of independent trials. On a single CPU core, these trials execute sequentially, meaning a simulation requiring 10 million paths might take minutes or hours, stalling the iterative design process.
The solution for these "embarrassingly parallel" workloads is the parfor loop provided by the Parallel Computing Toolbox. By distributing loop iterations across multiple CPU cores (workers), you can reduce execution time from minutes to seconds with minimal changes to your existing codebase.
How parfor Distributes Work
A standard for loop executes iterations one by one. In contrast, a parfor loop partitions the total number of iterations into chunks and assigns those chunks to separate workers in a parallel pool.
To use this effectively, MATLAB must categorize the variables inside the loop. If a variable is sliced, each worker receives only the specific index of the array it needs. If a variable is a reduction variable (like a running sum), MATLAB handles the communication to combine the results from all workers once the loop completes. If a variable is neither, it is treated as a broadcast variable, meaning a full copy is sent to every worker, which can consume significant memory.
Implementation: Pricing European Call Options
Consider a scenario where we need to price a European call option by simulating one million possible price paths. This requires calculating the payoff for each path and averaging them.
The Implementation Script
% Setup parameters
S0 = 100; % Initial stock price
K = 105; % Strike price
T = 1; % Time to maturity (years)
r = 0.05; % Risk-free rate
sigma = 0.2; % Volatility
numPaths = 1e6;% Number of simulations
% 1. Initialize the parallel pool (if not already running)
% This command should be run in the MATLAB Command Window or script
% Permissions: Requires Parallel Computing Toolbox license
if isempty(gcp('nocreate'))
parpool('local');
end
% Preallocate results array
payoffs = zeros(numPaths, 1);
% 2. Parallel Execution
tic;
parfor i = 1:numPaths
% Ensure independent random streams per worker
% MATLAB's rand() handles this automatically in parfor since R2015b
z = randn();
ST = S0 * exp((r - 0.5 * sigma^2) * T + sigma * sqrt(T) * z);
payoffs(i) = max(ST - K, 0);
end
parallelTime = toc;
% Calculate final price
optionPrice = exp(-r * T) * mean(payoffs);
fprintf('Parallel Execution Time: %.4f seconds\n', parallelTime);
Verification and Diagnostics
To verify the performance gain, run the same logic using a standard for loop. On a typical 4-core machine, you should see a speed-up approaching 4x, minus the overhead of managing the workers. You can check the active worker count by running numworkers(gcp) in the Command Window.
Critical Constraints and Trade-offs
Parallelization is not a "free" performance boost. There are three primary limitations to consider:
- Pool Startup Overhead: Launching a
parpooltakes several seconds. If your total simulation time is only 2 seconds, the overhead of starting the pool will make the parallel version slower than the serial one. Always keep the pool open across multiple simulation runs. - Memory Exhaustion: If the loop body creates large temporary arrays that are not sliced, each worker will allocate its own copy of that memory. On a machine with 16GB of RAM and 8 workers, a 3GB temporary array will crash the session.
- Random Number Correlation: In older versions of MATLAB, workers might share the same random seed, leading to identical paths. In R2015b and later,
parforautomatically initializes independent streams, but for specialized requirements, useRandStream.createto explicitly define the seed per worker.
Actionable Summary
Use parfor when your loop iterations are independent and the computational cost of the loop body outweighs the overhead of data distribution. To optimize, ensure your output arrays are sliced and avoid creating large, non-sliced variables inside the loop. If you notice a lack of speed-up, use the Parallel Monitor to check if workers are idling or if the system is bottlenecked by memory swapping.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.