Lazy‑mapped UMat versus explicit map/unmap for deterministic GPU‑CPU synchronization
The goal is to reduce host‑GPU synchronization latency in OpenCV pipelines while ensuring that GPU buffers allocated as UMat are never inadvertently exposed to host code. Lazy‑mapped UMat (the default) postpones the actual transfer until the data is first read on the host, which can avoid unnecessary copies but introduces variable stalls when the first acces