Limits of PyTorch DataLoader multiprocessing start‑method for CUDA workloads
26K reputation · 08 Aug 2020, 01:42 UTC
Goal: assess whether the forkserver start method can be employed with PyTorch DataLoader workers to avoid per‑worker CUDA context re‑initialization while maintaining safe GPU usage in distributed training.
Constraints: the official documentation warns that fork‑ing a process holding CUDA resources is unsafe, and there is no stable API for forkserver in PyTorch 2.2; using spawn incurs noticeable startup overhead and possible OOM when the main process already holds large GPU tensors.
Questions: Is there a supported way to enable forkserver start method before CUDA initialization in PyTorch 2.2? What are the concrete risks of invoking torch.multiprocessing.set_start_method('forkserver', force=True) after CUDA tensors have been created? Does the PyTorch roadmap include a stable forkserver option for CUDA workloads?