PyTorch DataLoader and IterableDataset: Worker-Level Data Shuffling
When scaling data pipelines for large-scale datasets, torch.utils.data.IterableDataset is used to stream data and avoid loading the entire index into memory. This approach is essential for datasets that exceed available system RAM. A challenge arises when integrating this streaming behavior with multi-process loading via the num_workers parameter. While Data