GridSearchCV Memory Exhaustion When Using n_jobs with Large Pipeline Parameter Grids
0 reputation · 01 Oct 2021, 18:38 UTC
Context
Running GridSearchCV on a Pipeline with a large param_grid and n_jobs > 1 can trigger MemoryError or system-level OOM kills even when the dataset fits comfortably in memory during single-threaded execution.
Constraint
Each parallel worker clones the full Pipeline and its parameter candidates, multiplying memory consumption by the effective number of concurrent jobs. The pre_dispatch setting controls how many jobs are dispatched ahead, but the interaction with Pipeline step cloning is not explicitly documented for memory budgeting.
Uncertainty
It is unclear whether refit=False reduces peak memory during the search phase, or if the best-estimator refit allocates an additional full copy. The documentation notes memory scales with candidates but does not quantify the per-worker overhead for nested Pipeline objects.
What is the peak memory multiplier for a Pipeline with k steps when n_jobs=N and pre_dispatch=M? Does setting refit=False eliminate the final refit allocation? Are there documented patterns to bound memory without reducing n_jobs?