Speed perturbation versus extended epochs for low-resource Kaldi nnet3 chain training under GPU memory constraints
0 reputation · 27 Dec 2022, 08:33 UTC
Goal: decide whether to apply speed perturbation or to increase training epochs when training a low-resource Kaldi nnet3 chain model under a fixed GPU memory budget.
Constraint: speed perturbation triples the effective training set, raising storage demand and I/O pressure while reducing the number of epochs needed for convergence; extended epochs keep the original data size, preserving lower memory footprint but increase GPU time and risk of overfitting on small corpora.
Uncertainty: the interaction between data augmentation, epoch count, GPU memory, and storage bandwidth is not analytically predictable, and the optimal balance varies with corpus size, model architecture and hardware.
Questions:
- For a given GPU memory limit, which strategy - speed perturbation or extra epochs - produces the lowest word error rate?
- How does the additional I/O and storage overhead of speed perturbation affect overall training throughput compared with extra epochs?
- At what corpus size does extended epoch training begin to overfit noticeably relative to speed-perturbed training?