The Hard Truth: No In-Place Retry for NCCL Timeouts
When NCCL times out during DDP gradient synchronization, PyTorch tears down the entire process group. This is by design—not a bug. The collective operations (all-reduce, all-gather) are atomic across all ranks, and a timeout on any single GPU means the operation cannot be considered complete on any rank.
Why Manual Retries Fail
Attempting to catch the RuntimeError and call dist.all_reduce again is unsafe because:
- The process group state is already destroyed—subsequent collectives will hang or fail immediately
- Even if the group were intact, ranks that did not timeout would have already applied gradients, causing weight divergence
- Optimizer step state would be inconsistent across ranks
The Only Safe Recovery Path
Use torch.distributed.checkpoint (PyTorch 2.3+) to save sharded checkpoints, then restart from the last known good state. The workflow is:
- Checkpoint regularly: Save model, optimizer, and RNG state every N steps
- Enable elastic restart: Use
torchrun --rdzv_backend=c10d --max_restarts=3 with TorchElastic
- Handle the exception: Catch the timeout, let TorchElastic restart workers
- Verify continuity: Ensure loss curves resume from checkpointed state
Model Consistency: All or Nothing
PyTorch guarantees consistency only when all ranks either complete the collective successfully or all roll back to the same checkpoint. There is no documented mechanism for partial recovery. If one rank recovers while others restart, parameters will diverge silently—this is not a recovery, it's corruption.
Diagnosis Checklist
Before increasing NCCL_TIMEOUT, investigate the root cause:
- Check GPU utilization: OOM can cause one rank to hang during backward
- Verify network health: InfiniBand or Ethernet congestion
- Review NCCL logs with
NCCL_DEBUG=INFO and NCCL_DEBUG_SUBSYS=ALL
- Confirm batch size consistency: Uneven batches can cause rank imbalance
Key Parameters
| Parameter | Value | Notes |
NCCL_TIMEOUT | 30-60 minutes | In milliseconds; longer values delay failure detection |
NCCL_BLOCKING_WAIT | 1 | Enables blocking waits for debugging |
find_unused_parameters | False | True adds extra collectives, increases timeout risk |
Bottom Line
For NCCL timeout recovery, treat it as a process-level failure, not an operation-level one. The only recommended approach is checkpoint-based restart with TorchElastic. Any attempt at in-place retry risks silent model corruption across your distributed training run.