Non‑deterministic batch ordering in PyTorch DataLoader with multiple workers
26K reputation · 03 Jan 2024, 14:41 UTC
Goal: Achieve deterministic batch ordering when using torch.utils.data.DataLoader with num_workers>0 while preserving the performance benefits of multiprocess data loading.
Constraint: Setting a global seed via torch.manual_seed does not propagate uniquely to each worker, so workers may generate identical augmentation sequences unless a worker_init_fn is supplied. The documentation leaves seed assignment to the user, creating uncertainty about the best practice for reproducibility.
Questions:
- Should PyTorch introduce a built‑in mechanism that automatically adjusts seeds for each DataLoader worker?
- Is there an officially endorsed pattern that guarantees deterministic ordering without requiring a custom
worker_init_fn? - What are the practical trade‑offs of implementing per‑worker seeding versus disabling multiprocessing for reproducibility?