Answer to the primary question
PyTorch does not automatically adjust seeds for each DataLoader worker. The standard, officially endorsed pattern for deterministic ordering is to combine a fixed torch.Generator for the sampler with a worker_init_fn that seeds each worker based on its worker_id. This keeps the multiprocessing performance while guaranteeing reproducible batch order.
Why this works
The sampler (e.g., RandomSampler) draws indices from a pseudorandom generator. By passing the same torch.Generator instance to every worker, you ensure the index sequence is the same across runs. The worker_init_fn then re‑initializes the global RNG state in each worker so that any augmentation code that relies on torch.manual_seed, random, or numpy.random behaves consistently.
Practical steps for your setup
torch.manual_seed(42) in the main process to set a baseline seed.
- Instantiate a
torch.Generator and seed it with the same value:
gen = torch.Generator()
gen.manual_seed(42)
- Create the sampler with this generator:
sampler = torch.utils.data.RandomSampler(dataset, generator=gen)
- Define a
worker_init_fn that seeds each worker uniquely:
def worker_init_fn(worker_id):
# Use worker_id to derive a distinct seed
seed = 42 + worker_id
torch.manual_seed(seed)
# If you use numpy or random, seed them too
import numpy as np, random
np.random.seed(seed)
random.seed(seed)
- Pass the sampler and
worker_init_fn to the DataLoader:
dataloader = torch.utils.data.DataLoader(
dataset,
batch_size=32,
sampler=sampler,
num_workers=4,
worker_init_fn=worker_init_fn
)
- Run the training loop. The first few batches should be identical across separate runs with the same seed.
Trade‑offs
- Per‑worker seeding preserves the parallelism benefits of
num_workers>0 and yields deterministic ordering, but you must add the worker_init_fn boilerplate.
- Disabling multiprocessing (setting
num_workers=0) eliminates the need for seeding logic but incurs a noticeable slow‑down because data loading becomes synchronous with the training loop.
- Both approaches still require that any CUDA operations that are non‑deterministic (e.g., CuDNN algorithms) are disabled if full reproducibility is desired.
Missing Diagnostic Detail
To fine‑tune the recommendation, could you confirm whether you are using a RandomSampler or a custom sampler that introduces randomness? This determines whether the generator passed to the sampler is sufficient or if additional seeding logic is needed inside the dataset.