Preventing Data Leakage with scikit‑learn Pipeline and ColumnTransformer
Learn how to wrap preprocessing steps in a Pipeline so that statistics are learned only on training data, avoiding leakage when modeling mixed numeric and categorical features.
22 Jul 2026, 07:48 UTC

Problem: preprocessing mixed data without leaking information
When a dataset contains both numeric columns (e.g., square footage, year built) and categorical columns (e.g., neighborhood, roof type), each type usually needs its own preprocessing: numeric values may require imputation and scaling, while categorical values need encoding. If you compute statistics such as means or encodings on the full dataset before splitting into training and test sets, information from the test set leaks into the model. This leads to overly optimistic performance estimates and can cause overfitting.
Thesis: a single Pipeline with ColumnTransformer isolates fitting to the training fold
By placing preprocessing steps inside a scikit-learn Pipeline that uses a ColumnTransformer, the fitting of imputers, scalers, and encoders happens only on the training data during fit. The same learned parameters are then applied to validation or test data via transform. The result is a single reusable object that can be used for training, hyperparameter tuning, and inference without risking leakage.
Section 1: Building the preprocessing blocks
First define separate pipelines for numeric and categorical columns.
from sklearn.pipeline import Pipeline
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import StandardScaler, OneHotEncoder
numeric_features = ['area', 'age', 'bathrooms']
categorical_features = ['neighborhood', 'roof_type']
numeric_transformer = Pipeline(steps=[
('imputer', SimpleImputer(strategy='median')),
('scaler', StandardScaler())
])
categorical_transformer = Pipeline(steps=[
('imputer', SimpleImputer(strategy='most_frequent')),
('onehot', OneHotEncoder(handle_unknown='ignore'))
])
The numeric pipeline fills missing values with the median and then scales to zero mean and unit variance. The categorical pipeline fills missing values with the most frequent category and then applies a one‑hot encoding that ignores unknown categories during transformation.
Section 2: Combining with ColumnTransformer and chaining to a model
Next, combine the two transformers with a ColumnTransformer and place the result inside a final Pipeline together with a model, for example a GradientBoostingRegressor.
from sklearn.compose import ColumnTransformer
from sklearn.ensemble import GradientBoostingRegressor
preprocess = ColumnTransformer(
transformers=[
('num', numeric_transformer, numeric_features),
('cat', categorical_transformer, categorical_features)
])
model = GradientBoostingRegressor(random_state=42)
clf = Pipeline(steps=[
('preprocess', preprocess),
('regressor', model)
])
At this point clf is a single object that knows how to preprocess data and then predict.
Section 3: Worked example with the Ames Housing dataset
The following snippet loads the Ames Housing dataset, splits it, fits the pipeline, and evaluates the root‑mean‑squared error (RMSE) on a hold‑out test set.
import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_squared_error
# Load data (requires pandas and the openml version of Ames Housing)
df = pd.read_csv('https://raw.githubusercontent.com/ageron/handson-ml2/master/datasets/housing/housing.csv')
# For illustration we keep only a subset of columns
cols = ['area', 'age', 'bathrooms', 'neighborhood', 'roof_type', 'SalePrice']
df = df[cols].dropna(subset=['SalePrice'])
X = df.drop('SalePrice', axis=1)
y = df['SalePrice']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=1)
# Fit the pipeline
clf.fit(X_train, y_train)
# Predict and compute RMSE
y_pred = clf.predict(X_test)
rmse = mean_squared_error(y_test, y_pred, squared=False)
print(f'Test RMSE: {rmse:.2f}')
Running the example should produce a finite RMSE value without any errors, assuming scikit-learn>=1.0, pandas, and numpy are installed.
Trade‑off/limitation: high‑cardinality categorical features
OneHotEncoder creates a new column for each distinct category. When a feature has many unique values (e.g., ZIP codes or product IDs), the resulting matrix can become very large, increasing memory usage and training time, especially for linear models that benefit from sparsity. Two common mitigations are:
- Target encoding: replace each category with the mean of the target for that category. This reduces dimensionality but can introduce leakage if not computed within each cross‑validation fold.
- Feature hashing (
FeatureHasher): maps categories to a fixed‑size vector via a hash function, trading a small risk of collisions for bounded dimensionality.
Both alternatives require careful validation to ensure they do not degrade model performance.
Verification steps
To confirm that the pipeline is protecting against leakage:
- Compare imputation statistics: compute the median of a numeric column on the full dataset and on the training split only. The median learned by the pipeline’s
SimpleImputer(accessible viaclf.named_steps['preprocess'].named_transformers_['num'].named_steps['imputer'].statistics_) should match the training‑only median, not the full‑dataset median. - Reproduce predictions manually: after fitting, call
clf.predict(X_test)and separately apply the fittedColumnTransformertoX_testthen call the model’spredict. The two outputs should be identical (within floating‑point tolerance). - Check output format: verify that
clf.predictreturns a dense NumPy array (tree‑based models require dense input). If you switch to a linear model that accepts sparse input, ensure theColumnTransformer is configured to output sparse matrices where possible.
Actionable closing
Wrap all preprocessing steps that depend on data statistics inside a Pipeline that uses a ColumnTransformer. This guarantees that fitting occurs only on the training data, eliminates a common source of leakage, and gives you a single object to pass to grid search, cross‑validation, or production inference. When dealing with high‑cardinality categorical features, consider target encoding or hashing, but always validate that the alternative does not introduce bias or leakage.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.