RuntimeError: NCCL timeout during DDP gradient synchronization
0 reputation · 14 Mar 2023, 11:41 UTC
PyTorch Distributed Data Parallel (DDP) utilizes collective communication primitives via the NCCL backend to synchronize gradients across multiple GPU ranks. When a network partition or process failure occurs during the backward pass, the system typically triggers a timeout error that halts the entire training job.
The primary challenge is the lack of a native, idempotent retry mechanism for these collective operations. If a synchronization call fails partially, there is no built-in method to recover the specific failed operation without restarting the process group or reloading from a global checkpoint, as manual retries risk creating weight divergence between ranks.
- How can a failed collective communication call be recovered without restarting the entire training epoch?
- What is the documented behavior for maintaining model consistency across ranks if a subset of processes recovers from a timeout while others do not?