Fortran coarray collectives: non-coarray write visibility across images without sync memory
26.6K reputation · 26 Nov 2022, 14:20 UTC
Symptom
Portable reasoning about data visibility after a collective subroutine (co_sum, co_max, co_min, co_reduce) is hindered because the Fortran 2018 standard specifies an implicit synchronization but does not mandate whether that synchronization acts as a full memory fence equivalent to sync memory.
Context
Current compilers diverge: GNU Fortran 12+ treats collectives as sequentially consistent barriers, while Intel Fortran (ifx) 2023+ and NVIDIA HPC SDK 23.11 document them as image-control statements with acquire/release semantics only. The standard's term "synchronization" in clause 11.7.2 is ambiguous regarding memory-model strength, and no conformance test in the Fortran Standard Test Suite validates the ordering guarantee.
Goal
Determine whether a non-coarray write performed by image 1 before calling a collective subroutine is guaranteed to be visible to image 2 after its matching collective returns, without an intervening sync memory.
Specific questions:
- Does the Fortran 2023 standard or any published interpretation clarify the memory-ordering strength of collective subroutines?
- Is there a compiler-agnostic test pattern that reliably exposes the visibility difference between the acquire/release and seq-cst models?
- For mixed-language programs using ISO_C_BINDING, can C11
atomic_thread_fencebe used to order around Fortran collectives?