Memory limit constraints for n_jobs parallelization in GridSearchCV
26.5K reputation · 08 May 2022, 20:34 UTC
The n_jobs parameter in estimators like GridSearchCV and RandomForestClassifier uses the joblib backend to distribute tasks across CPU cores. Scikit-learn attempts to use memory mapping for large arrays to reduce duplication, but effectiveness depends on data format and process creation method.
In resource-constrained settings, n_jobs=-1 can increase memory pressure because there is no native configuration to cap memory per worker or to fall back to sequential execution when usage grows. Behavior differs between fork and spawn start methods, and memory mapping may not apply to non-contiguous arrays or object dtype inputs.
Is there a documented way to limit the memory footprint of individual joblib workers used by scikit-learn? Can memory mapping be guaranteed for non-contiguous NumPy arrays, and is there a supported signal or option to throttle parallelism based on memory usage?
1 answer
1 question comment
Use comments to ask for clarification. Post a solution as an answer.
26,525 reputation · 09 May 2022, 01:58 UTC
Besides limiting n_jobs, you can curb peak memory by adjusting pre_dispatch (e.g., pre_dispatch=2*n_jobs) so that only a subset of parameter fits are spawned at once, and by setting joblib.Parallel(..., max_nbytes=1e6) to force large arrays to be shared via memory‑mapped files rather than pickled copies. These controls work with the default loky backend and are effective when the input is a contiguous, numeric NumPy array.