MemoryError occurs when Pipeline with GridSearchCV uses n_jobs=-1 on large datasets
0 reputation · 19 Nov 2020, 06:06 UTC
Goal: Determine which transformer or estimator step dominates runtime in a scikit‑learn Pipeline before investing effort in hyperparameter tuning with GridSearchCV.
Constraint: scikit‑learn does not expose built‑in timers for individual Pipeline steps, and adding custom timing code may interfere with cross‑validation folds or parallel execution when n_jobs is set.
Uncertainty: Whether an external profiler can be safely applied inside a GridSearchCV loop without altering the fitted model or causing data leakage, and if the n_jobs parameter obscures per‑step measurements.
Questions: How can per‑step fit/transform times be measured reliably inside GridSearchCV? Does wrapping steps with a timing decorator introduce bias or leakage? Can step‑level timing be obtained when n_jobs > 1?