Using Fortran DO CONCURRENT for Safe Parallel Loops
Learn how to express independent loop iterations in modern Fortran with DO CONCURRENT, see a matrix‑vector product example, and understand the trade‑offs and verification steps needed for reliable parallel execution.
01 Aug 2026, 16:50 UTC

The problem: needing safe, compiler‑driven parallelism
Many scientific codes spend most of their time in loops that operate on independent data, such as updating each row of a matrix or each element of a vector. Writing explicit OpenMP or MPI directives can be verbose and error‑prone, especially when the goal is simply to tell the compiler that iterations do not depend on one another. Fortran’s DO CONCURRENT construct addresses this by letting you declare independence directly in the loop syntax, enabling the compiler to parallelize, reorder, or vectorize the loop when it is safe to do so.
How DO CONCURRENT works
The construct looks like a regular DO loop but carries additional semantics:
DO CONCURRENT (i=1:n, j=1:m)
! loop body
END DO
Key restrictions ensure safety:
- No data dependencies between different iterations (i.e., each iteration must read/write only private or explicitly shared data).
- No I/O statements,
EXIT, orCYCLEthat depend on the loop indices. - All variables used inside must be either
LOCAL(private to each iteration) or explicitly given a sharing attribute (SHARED,DEFAULT(NONE)with explicit specification).
When these conditions hold, the compiler may:
- Execute iterations in parallel threads (if OpenMP support is enabled).
- Reorder iterations for better cache usage.
- Vectorize the loop using SIMD instructions.
If the compiler cannot guarantee safety, it will fall back to sequential execution.
Worked example: matrix‑vector product
Consider computing y = A * x where A is an n × m matrix and x and y are vectors. Each row of A can be processed independently, making the outer loop a good candidate for DO CONCURRENT.
program matvec
implicit none
integer, parameter :: n = 1000, m = 1000
real, dimension(n,m) :: A
real, dimension(m) :: x
real, dimension(n) :: y, y_seq
integer :: i, j
! Initialize data (simple deterministic values)
A = reshape([(real(i*j, kind=4), i=1,n*m)], [n,m])
x = [(real(i, kind=4), i=1,m)]
y = 0.0
y_seq = 0.0
! Parallel version using DO CONCURRENT
DO CONCURRENT (i=1:n)
y(i) = 0.0
DO j=1,m
y(i) = y(i) + A(i,j) * x(j)
END DO
END DO
! Sequential reference
DO i=1,n
y_seq(i) = 0.0
DO j=1,m
y_seq(i) = y_seq(i) + A(i,j) * x(j)
END DO
END DO
! Simple verification
if (maxval(abs(y - y_seq)) < 1.0e-6) then
print *, 'Results match within tolerance.'
else
print *, 'Mismatch detected!'
end if
end program matvec
Explanation:
- The outer loop over rows is marked
DO CONCURRENT. Inside, each iteration writes only to its own elementy(i)and reads from the shared, read‑only arraysAandx. No cross‑iteration dependencies exist. - The inner loop over columns remains sequential because it works on private scalars.
- After the parallel loop, a sequential reference computes
y_seqfor comparison.
To enable parallel execution with GNU Fortran, compile with optimization and OpenMP support:
gfortran -O3 -fopenmp -march=native matvec.f90 -o matvec
When run on a multi‑core CPU, the program should print "Results match within tolerance." and the parallel region will typically show reduced wall‑clock time compared with a purely sequential build (-O3 without -fopenmp).
Trade‑offs and limitations
While DO CONCURRENT can deliver significant speed‑ups for embarrassingly parallel workloads, there are practical considerations:
- Hidden dependencies: If a variable is inadvertently shared (e.g., through a pointer or a global variable) and written in more than one iteration, a race condition occurs, leading to nondeterministic results. Careful inspection of data flow is required.
- Compiler support: Older Fortran compilers (pre‑Fortran 2008) ignore the construct and treat it as a sequential loop. Even among modern compilers, the degree of parallelization varies; some may only vectorize, others may generate OpenMP threads. Consult the documentation for
gfortran,ifort, orflangto see what flags are needed. - Restricted statements: The prohibition on
EXIT,CYCLE, and I/O inside the concurrent loop can require refactoring existing code.
These limitations mean that DO CONCURRENT is best suited for loops where independence is obvious and where the codebase can be adjusted to meet the syntactic constraints.
Actionable steps for using DO CONCURRENT safely
- Identify a loop whose iterations read only from read‑only data and write to distinct locations.
- Mark the loop with
DO CONCURRENTand ensure all variables are eitherLOCALor explicitly declaredSHAREDwith proper intent. - Compile with optimization and the appropriate parallel flag (e.g.,
-O3 -fopenmpfor gfortran). - Verify correctness by comparing the output against a known‑good sequential version or by using a race‑detection tool such as ThreadSanitizer (if supported).
- Measure performance with a simple timing intrinsic like
CPU_TIMEand confirm scaling with the number of available cores.
Following these steps lets you harness the compiler’s ability to parallelize safe loops without the verbosity of explicit threading APIs, while keeping the risk of incorrect results under control.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.