Local Memory vs Global Memory Tiling: Which Wins on GPUs with Tiny Local Memory?
When optimizing a matrix‑multiplication kernel for an OpenCL device, the first step is to quantify the memory‑traffic bottleneck. The decision is whether to tile the work‑group using local memory or to rely on global memory accesses, especially on GPUs whose CL_DEVICE_LOCAL_MEM_SIZE is only a few kilobytes. Local memory can dramatically cut global traffic, b