Building Reproducible Models with scikit-learn Pipeline and ColumnTransformer
Learn how to combine numeric and categorical preprocessing into a single Pipeline to avoid data leakage and simplify model training.
03 Feb 2026, 20:19 UTC

The problem: preprocessing mixed data without leaking information
When a dataset contains both numeric columns (e.g., age, salary) and categorical columns (e.g., city, department), a typical workflow is to scale the numeric features and one‑hot encode the categorical ones, then concatenate the results before fitting a model. Doing these steps manually outside of the model object creates two risks:
- During cross‑validation or hyperparameter search, the preprocessing statistics (like mean and variance) can be computed on the whole fold, leaking information from the validation set into the training set.
- Any change in the preprocessing code must be mirrored in every place the model is used (training, testing, production), increasing the chance of inconsistency.
scikit-learn’s Pipeline and ColumnTransformer solve both issues by bundling preprocessing and estimation into a single, reproducible object.
How ColumnTransformer works
ColumnTransformer lets you specify a list of transformers, each paired with the column names (or indices) it should act on. For example:
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import StandardScaler, OneHotEncoder
numeric_features = ['age', 'salary', 'years_experience']
categorical_features = ['city', 'department', 'education_level']
preprocess = ColumnTransformer(
transformers=[
('num', StandardScaler(), numeric_features),
('cat', OneHotEncoder(handle_unknown='ignore'), categorical_features)
])
The transformer applies StandardScaler only to the numeric columns and OneHotEncoder only to the categorical columns, then horizontally stacks the results. No manual np.hstack or pd.concat is required.
Wrapping everything in a Pipeline
Once the preprocessing step is defined, you can place it at the start of a Pipeline followed by any estimator that implements fit and predict:
from sklearn.pipeline import Pipeline
from sklearn.linear_model import LogisticRegression
model = Pipeline(steps=[
('preprocess', preprocess),
('classifier', LogisticRegression(max_iter=1000))
])
When you call model.fit(X_train, y_train), the pipeline:
- Fits each transformer on the training data only.
- Transforms the training data using those fitted transformers.
- Fits the estimator on the transformed data.
During model.predict(X_test) the same fitted transformers are applied to the test data, guaranteeing that no information from the test set influences the training phase.
Worked example
Assume a CSV file employees.csv with the columns described above and a binary target promoted. The following script loads the data, splits it, builds the pipeline, fits it, and prints the hold‑out accuracy.
import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score
# Load data
df = pd.read_csv('employees.csv')
X = df.drop('promoted', axis=1)
y = df['promoted']
# Train‑test split
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42, stratify=y
)
# Define preprocessing (same as above)
numeric_features = ['age', 'salary', 'years_experience']
categorical_features = ['city', 'department', 'education_level']
preprocess = ColumnTransformer(
transformers=[
('num', StandardScaler(), numeric_features),
('cat', OneHotEncoder(handle_unknown='ignore'), categorical_features)
])
# Build pipeline
model = Pipeline(steps=[
('preprocess', preprocess),
('classifier', LogisticRegression(max_iter=1000))
])
# Fit
model.fit(X_train, y_train)
# Predict and evaluate
y_pred = model.predict(X_test)
print('Hold‑out accuracy:', accuracy_score(y_test, y_pred))
Run this script in a fresh Python environment where scikit‑learn ≥ 0.20 is installed (e.g., pip install scikit-learn pandas). No special permissions are needed; the script only reads the local CSV file.
Trade‑offs and limitations
While the pipeline approach eliminates leakage and simplifies code, there are practical considerations:
- Memory usage:
OneHotEncodercreates a sparse matrix by default, but if you later convert to dense (e.g., by passing the output to an estimator that expects dense arrays) memory can grow quickly with high‑cardinality features. In such cases consider usingTargetEncoderfromcategory_encodersor hashing tricks. - API changes: Between scikit‑learn 0.22 and 0.24 the handling of sparse output from
ColumnTransformerwas adjusted. If you upgrade, re‑run the verification steps below to ensure the pipeline still behaves as expected. - Debugging: Because the preprocessing steps are hidden inside the pipeline, inspecting intermediate arrays requires accessing the named step, e.g.,
model.named_steps['preprocess'].transform(X_test).
How to verify the pipeline works correctly
After fitting, you can perform three quick checks:
- Output shape: Confirm that
model.predict(X_test)returns a 1‑D array with length equal to the number of test samples. - Feature count: Inspect the number of features after preprocessing:
n_features = model.named_steps['preprocess'].transform(X_test).shape[1] print('Number of processed features:', n_features) - Numerical equivalence: Manually preprocess a copy of the training data using the same transformers, fit a fresh
LogisticRegressionon the manually processed data, and compare its coefficients (or predictions) with those from the pipeline. They should match up to floating‑point tolerance.
If any of these checks fail, revisit the column lists or transformer parameters.
Actionable closing
Start every new modeling project with a Pipeline that begins with a ColumnTransformer tailored to your data types. This habit guarantees reproducible preprocessing, reduces boilerplate code, and protects you from subtle data‑leakage bugs. When you encounter high‑cardinality categorical variables, evaluate alternative encoders before defaulting to one‑hot encoding, and always re‑run the verification steps after upgrading scikit‑learn.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.