Parallelizing Julia for Loops with @threads: A Practical Guide
Learn how to accelerate CPU‑bound Julia loops using the @threads macro. This guide covers setting the thread count, writing race‑free code, verifying thread usage, and troubleshooting common pitfalls.
31 Mar 2026, 15:05 UTC

Desired Outcome
Speed up a CPU‑bound for loop in Julia by distributing its iterations across all available CPU cores using the @threads macro, while keeping the code safe from data races.
Prerequisites
- Julia 1.6 or newer (threading is enabled by default).
Base.Threads.nthreads()should return >1. - Understanding of basic Julia syntax and the
forloop construct. - Knowledge that shared mutable state must be protected in a multithreaded context.
Setting the Thread Count
Julia determines the number of active threads at startup. Set the environment variable JULIA_NUM_THREADS before launching the interpreter:
# On Linux/macOS
export JULIA_NUM_THREADS=8
julia
# On Windows (PowerShell)
$env:JULIA_NUM_THREADS = 8
julia
Inside Julia you can verify the count with:
using Base.Threads
println("Threads: ", nthreads())
Changing JULIA_NUM_THREADS after the REPL has started has no effect.
Writing a Race‑Free @threads Loop
Only loop bodies that do not modify shared mutable variables can safely use @threads. If you need to accumulate results, use thread‑local storage or atomic operations.
Example: Sum of Squares (Thread‑Local Accumulation)
using Base.Threads
function sum_of_squares(n::Int)
local_sums = zeros(Float64, nthreads())
@threads for i in 1:n
local_sums[threadid()] += i^2
end
return sum(local_sums)
end
# Serial version for comparison
function sum_of_squares_serial(n::Int)
s = 0.0
for i in 1:n
s += i^2
end
return s
end
n = 10^7
@time sum_of_squares_serial(n) # baseline
@time sum_of_squares(n) # threaded
In this pattern each thread writes only to its own slot in local_sums, eliminating data races. The final sum combines the partial results.
Common Pitfall: Shared Mutable Counter
counter = 0
@threads for i in 1:10^7
counter += 1 # Data race!
end
Running this will often produce a value less than 10^7 because increments overlap. Use Threads.Atomic or a lock:
using Base.Threads
counter = Threads.Atomic{Int}(0)
@threads for i in 1:10^7
Threads.atomic_add!(counter, 1)
end
println(counter[]) # Correct result
Expected Checks
- Confirm thread count:
println(nthreads()). - Verify distribution: insert
@show threadid()inside the loop; output should show a spread across 1…nthreads(). - Compare runtimes:
@timeorBenchmarkTools.@btimefor serial vs threaded loops. - Validate correctness: ensure the result matches the serial version.
Recovery Options
- If a package is not thread‑safe, run it outside the
@threadsblock or wrap it in aThreads.@spawnwith a dedicated worker. - When encountering data races, switch to thread‑local accumulation or use
Threads.Atomicas shown. - For very short iterations, the overhead of thread scheduling may outweigh benefits. Profile with
BenchmarkToolsto decide whether to parallelize. - To change thread count after a session, exit Julia and restart with a new
JULIA_NUM_THREADSvalue.
Limitations and Practical Checks
• The @threads macro only works on simple for loops; nested loops or complex comprehensions are not automatically parallelized.
• The number of threads is fixed at startup; dynamic scaling is not supported.
• Some third‑party packages internally use global state that is not thread‑safe. If your loop calls such functions, you may need to serialize those calls.
• Always run a small benchmark before scaling to large data sets. A quick @btime test can reveal whether parallelism yields a measurable speedup.
Conclusion
Using @threads in Julia can deliver significant speedups for CPU‑bound loops, provided you respect thread safety rules and verify the distribution of work. By setting JULIA_NUM_THREADS, protecting shared state, and benchmarking, you can confidently parallelize loops without sacrificing correctness.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.