Local Memory vs Global Memory Tiling: Which Wins on GPUs with Tiny Local Memory?
0 reputation · 02 Apr 2021, 06:06 UTC
When optimizing a matrix‑multiplication kernel for an OpenCL device, the first step is to quantify the memory‑traffic bottleneck. The decision is whether to tile the work‑group using local memory or to rely on global memory accesses, especially on GPUs whose CL_DEVICE_LOCAL_MEM_SIZE is only a few kilobytes.
Local memory can dramatically cut global traffic, but allocating a tile that exceeds the device’s local‑memory quota triggers ambiguous runtime behavior: some vendors silently truncate, others return CL_INVALID_VALUE, and a few allow execution with an undefined region. This spec gap makes it hard to predict performance gains.
Key constraints to evaluate are:
- Actual local memory size and how many work‑groups can share it.
- Allocation overhead and possible spill to global memory.
- Vendor‑specific handling of oversize local‑memory requests.
Given these uncertainties, how does an oversized local‑memory allocation behave on different GPUs? Does tiling with local memory still reduce global traffic on a device with only 4 KB of local memory? At what tile size does the allocation overhead outweigh the benefit of reduced global traffic?