GridSearchCV in scikit-learn: Getting the Scoring, Refit, and Parallelism Right
GridSearchCV's defaults are easy to misread. Here's how scoring, refit, and n_jobs really behave in scikit-learn, plus when to switch to RandomizedSearchCV or TimeSeriesSplit.
04 Dec 2025, 10:48 UTC

You've trained a random forest, the accuracy looks fine, and now someone asks the obvious question: did you tune it? GridSearchCV is the standard answer in scikit-learn, but the defaults are easy to misread. A search that silently optimizes the wrong metric, or one that returns a model you can't actually use, is worse than no search at all. The good news: three parameters — scoring, refit, and n_jobs — control almost everything that goes wrong, and they're simple once you know what each one really does.
What GridSearchCV actually does
GridSearchCV takes a dictionary of hyperparameter candidates and evaluates every combination using k-fold cross-validation. For each combination, the data is split into k folds; the model trains on k−1 folds and validates on the remaining one, rotating until every fold has served as validation. The average validation score estimates how well that combination generalizes. After scoring all combinations, the search picks a winner.
Two consequences follow. First, cost multiplies fast: 5 values each for 3 parameters with 5-fold CV means 625 model fits. Second, the "best" result is only as meaningful as the metric you scored it against — which is where the first trap lives.
Trap one: the default scoring metric
If you don't pass scoring, GridSearchCV uses the estimator's own score() method. For classifiers that's accuracy; for regressors, R². That sounds reasonable until you remember why you're tuning in the first place. On an imbalanced fraud dataset, accuracy happily rewards a model that predicts "legitimate" every time.
The fix is to state the metric explicitly, using either a string from scikit-learn's scoring registry or a callable built with make_scorer:
from sklearn.model_selection import GridSearchCV
from sklearn.ensemble import RandomForestClassifier
param_grid = {
"n_estimators": [100, 300],
"max_depth": [None, 10, 20],
"min_samples_leaf": [1, 4],
}
search = GridSearchCV(
RandomForestClassifier(random_state=42),
param_grid,
scoring="f1", # explicit: optimize F1, not accuracy
cv=5,
n_jobs=-1,
refit=True,
)
search.fit(X_train, y_train)
print(search.best_params_, search.best_score_)Run this in any Python environment with scikit-learn installed (the examples assume version 1.0 or newer — check with import sklearn; print(sklearn.__version__)). No special permissions are needed. If you need several metrics at once, pass a dict of scorers and set refit to the name of the metric that should decide the winner, e.g. refit="f1".
Trap two: forgetting what refit does
With refit=True (the default), GridSearchCV retrains the winning configuration on the entire training set after the search, and search.best_estimator_ is that fitted model ready for prediction. Set refit=False and you get the scores and best parameters, but best_estimator_ won't exist — calling search.predict() raises an error. That's a legitimate choice when you only want to compare configurations, but it surprises people who then try to deploy the result.
Also note the asymmetry: cross-validation scores come from models trained on 4/5 of the data, while the refit model trains on all of it. The best_score_ is therefore a slightly pessimistic estimate of the final model. For an honest evaluation, keep a held-out test set the search never touches and score the refit estimator on it once, at the end.
Trap three: n_jobs=-1 and memory
n_jobs=-1 parallelizes fits across all CPU cores, which is usually the right call. But each worker needs its own copy of the data it touches, so on a large dataset you can multiply memory usage by your core count and watch the machine swap or the workers get killed. If that happens, reduce n_jobs to a small number, or tune pre_dispatch (for example pre_dispatch="2*n_jobs") to limit how many tasks are queued at once. A quick sanity check: run the same grid with n_jobs=1 and n_jobs=-1 on a small subsample and confirm the wall-clock time actually drops and best_params_ matches.
When the grid itself is the problem
Grid search is exhaustive, which is both its strength and its ceiling. Three parameters with five options each is fine; six parameters with continuous ranges is not. When the space gets large, RandomizedSearchCV samples a fixed number of combinations and typically finds a near-optimal point in a fraction of the time — you trade completeness for a budget you control via n_iter.
There's a second, subtler limitation: ordinary k-fold CV assumes rows are independent and exchangeable. For time-ordered data, random folds leak the future into training, inflating scores. Pass cv=TimeSeriesSplit(n_splits=5) instead so each fold only validates on data that comes after its training window.
A practical checklist
- Always pass
scoringexplicitly — never rely on the estimator default. - Keep a held-out test set outside the search for one final, honest evaluation.
- Leave
refit=Trueunless you deliberately only want scores. - Start with
n_jobs=-1, but watch memory and fall back if workers die. - Use
RandomizedSearchCVwhen the full grid would take hours, andTimeSeriesSplitwhen rows have a time order.
GridSearchCV isn't magic; it's a loop with good bookkeeping. Set the metric you actually care about, understand what refit gives you, and respect your data's structure — and the "did you tune it?" question answers itself.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.