Answer
As of PyTorch 2.2 there is no officially stable API for the forkserver start method when using CUDA. You can attempt to set it before any CUDA initialization, but the call may fail or be unsupported on some platforms. The only method guaranteed to work safely with CUDA workers is spawn.
Confirmed facts
- PyTorch’s DataLoader supports three start methods:
fork, spawn, and forkserver.
fork is unsafe with CUDA because forking a process that holds CUDA resources can corrupt the driver state.
spawn creates a brand‑new process, guaranteeing a fresh CUDA context in each worker; it is the only method that is reliably safe for CUDA workloads.
forkserver behaves like spawn but starts a dedicated server process first. It is safe for CUDA where available, but it is not guaranteed to exist on all systems (e.g., Windows) and is not marked as a stable API in PyTorch 2.2.
- Calling
torch.multiprocessing.set_start_method after any multiprocessing activity (including after CUDA tensors have been created) raises a RuntimeError.
Likely explanation
If you invoke torch.multiprocessing.set_start_method('forkserver', force=True) after CUDA tensors exist in the parent process, the server process that forkserver creates will have been forked after CUDA initialization. Although the workers themselves are spawned from that server, the inherited CUDA context can lead to undefined behavior, driver crashes, or out‑of‑memory errors. The safest approach is to set the start method before any CUDA operation and before creating the DataLoader.
Steps to try forkserver (if you are on a Unix‑like system)
- Ensure no CUDA tensors or CUDA‑initializing code has run yet.
- Set the start method as early as possible, preferably before importing
torch or immediately after:
import torch
try:
torch.multiprocessing.set_start_method('forkserver', force=True)
except ValueError as e:
# Already set or not supported
print('Cannot set forkserver:', e)
- Create your DataLoader with
num_workers>0 (optionally pass multiprocessing_context=torch.multiprocessing.get_context('forkserver')).
- Verify the method in effect:
print(torch.multiprocessing.get_start_method()) # should show 'forkserver'
- Inside the worker (e.g., your dataset’s
__getitem__), you can safely create CUDA tensors; each worker will have its own context.
When to ask for a missing diagnostic detail
The recommendation changes if you are running on Windows, where forkserver is not available. In that case you must use spawn. Please confirm your operating system if you are unsure whether forkserver can be attempted.