Fork or Spawn: Which Start Method Should a CUDA DataLoader Pipeline Standardize On?
0 reputation · 06 Oct 2024, 13:26 UTC
GPU training with DataLoader(num_workers>0) carries a start-method decision PyTorch leaves to the user. Python's long-standing Linux default, fork, starts workers almost instantly and lets them inherit parent memory, including CUDA state. spawn isolates each worker in a fresh interpreter, re-importing the module and re-initializing CUDA per worker.
That default is not stable across environments: Linux has historically forked, Windows hosts spawn, many production setups pin spawn for isolation, and recent Python releases have moved the Linux default away from fork. PyTorch guidance discourages fork once CUDA is initialized, yet fork is what local development has long relied on. Behavior has also shifted across PyTorch 1.x and 2.x, so the answer is version-sensitive; this assumes a current PyTorch 2.x release and should be checked against the pinned Python and PyTorch versions.
The trade-off is unresolved: fork buys startup speed but demands discipline about CUDA initialization order; spawn buys isolation but adds per-worker re-initialization and import-time sensitivity. Standardizing on one method across laptop and production is appealing — but which?
Should a CUDA pipeline pin spawn everywhere and absorb the worker startup cost, or pin fork on Linux with strict initialization ordering? And should the active start method be asserted at startup in both environments, so divergence fails fast instead of surfacing mid-training?