Does LLVM's LoopVectorize pass re‑evaluate its cost model after profile‑guided optimization marks a function hot or cold?
0 reputation · 19 Sept 2023, 10:52 UTC
0 reputation · 19 Sept 2023, 10:52 UTC
LLVM’s LoopVectorize pass decides whether to vectorize a loop based on a cost model that evaluates trip count, memory stride, alignment, and target‑specific vector width. The model is consulted during the pass’s execution, which typically runs before profile‑guided optimization (PGO) marks functions as hot or cold. After PGO, the optimizer may reorder or re‑run other passes, but it is unclear whether LoopVectorize is invoked again with the updated profile information or whether it retains its original cost‑model decision.
This raises the question of whether the vectorizer’s decision is stable across PGO iterations or if a hot/cold designation should trigger a re‑evaluation of vectorization profitability.
Does the LoopVectorize pass re‑run after PGO to incorporate the updated hot/cold flags? If it does not, what mechanisms exist to trigger a re‑evaluation of vectorization decisions when profile data changes? Are there existing passes or flags that force the vectorizer to reconsider its cost model after PGO?
28775 reputation · 19 Sept 2023, 13:01 UTC
The LoopVectorize pass evaluates its cost model once per loop when it runs. It uses whatever profile data is available at that moment, including the hot/cold flag that PGO emits for a function. If PGO marks a function as hot or cold before the pass runs, the cost model will incorporate that information and may change the vectorization decision. The pass does not perform a second cost‑model calculation after a function has already been marked; it simply relies on the profile data that is present when it executes.
.profraw files during a test run.-fprofile-use, the LoopVectorize pass reads the profile data and uses the hot/cold flag in its trip‑count estimate..profraw) is the only way to see a different vectorization outcome.clang -fprofile-generate -O2 -S -emit-llvm your.c -o your.bc
./your_executable # run to produce .profraw files
llvm-profdata merge -o merged.profdata *.profraw
clang -fprofile-use=merged.profdata -mllvm -debug-only=loop-vectorize -O2 your.c -S -emit-llvm -o vectorized.bc
LoopVectorize: cost model for loop @0x...: tripCount=1234 hot=true
__attribute__((cold)) to a function or manipulate the profile to change its count, then repeat steps 1‑3. A different trip‑count estimate should appear in the log.
-fprofile-use and let the optimization pipeline run fresh.-O3 or -fno-slp-vectorize to force different ordering of passes, but this does not guarantee a re‑evaluation of the cost model for the same loop.
Could you tell me which LLVM version you are using? The interaction between PGO and LoopVectorize has evolved, and the exact pass ordering can differ between releases.
Use comments to ask for clarification. Post a solution as an answer.
No question comments on this page.