Streamlining Mixed‑Type Data Preprocessing with scikit‑learn Pipeline and ColumnTransformer
Learn how to combine scikit‑learn's Pipeline with ColumnTransformer to preprocess mixed numeric and categorical data safely and reproducibly.
30 Nov 2025, 01:20 UTC

Problem: repetitive preprocessing code leads to leakage risk
When a dataset contains both numeric and categorical columns, it is common to write separate scaling and encoding steps, fit them on the training set, and then manually apply the same transformations to validation or test data. This approach is error‑prone: forgetting to refit a scaler on the training split or applying an encoder fitted on the whole dataset can introduce data leakage and make reproducibility harder.
Thesis: wrapping preprocessing in a Pipeline with ColumnTransformer eliminates boilerplate and guarantees consistent transforms
By chaining a ColumnTransformer (which applies different transformers to selected columns) with an estimator inside a scikit‑learn Pipeline, you obtain a single estimator that learns from the training data and applies the exact same transformations during prediction. The pipeline can be serialized with pickle or joblib, making deployment straightforward.
How ColumnTransformer works
ColumnTransformer takes a list of tuples: (name, transformer, columns). Each transformer is fitted only on the columns it receives; the outputs are concatenated horizontally. This lets you apply, for example, StandardScaler to numeric columns and OneHotEncoder to categorical columns in a single step.
Worked example
import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score
# Load a CSV with mixed types
# df = pd.read_csv('employees.csv')
# Assume columns: age (numeric), salary (numeric), department (categorical), left (target)
# Define preprocessing for numeric and categorical features
numeric_features = ['age', 'salary']
numeric_transformer = StandardScaler()
categorical_features = ['department']
categorical_transformer = OneHotEncoder(handle_unknown='ignore')
preprocessor = ColumnTransformer(
transformers=[
('num', numeric_transformer, numeric_features),
('cat', categorical_transformer, categorical_features)
])
# Create the full pipeline
clf = Pipeline(steps=[
('preprocessor', preprocessor),
('classifier', LogisticRegression(max_iter=1000))
])
# Split data – fit only on training set
X = df.drop('left', axis=1)
y = df['left']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
# Fit the pipeline
clf.fit(X_train, y_train)
# Predict and evaluate
y_pred = clf.predict(X_test)
print('Accuracy:', accuracy_score(y_test, y_pred))
# Inspect the learned steps
print('Pipeline steps:', clf.named_steps.keys())
The example assumes scikit‑learn ≥ 0.20 (the version that introduced ColumnTransformer). If you run the code, you can verify that the pipeline’s named_steps contain both 'preprocessor' and 'classifier', and that the shape of clf.transform(X_test) matches the number of numeric columns plus the one‑hot encoded categories.
Trade‑off and limitation
The main drawback is increased debugging complexity: when an error occurs inside the ColumnTransformer, the traceback points to the transformer rather than your original code, and inspecting intermediate arrays requires accessing the pipeline’s steps. However, this complexity is offset by the elimination of manual fit/transform splits, which greatly reduces the chance of leakage and simplifies model export.
Actionable closing
Adopt the Pipeline + ColumnTransformer pattern whenever you need to preprocess heterogeneous features. Start by identifying numeric and categorical columns, assign appropriate transformers, and wrap everything in a Pipeline. Verify the installed scikit‑learn version (sklearn.__version__) and always fit the pipeline on training‑only data to maintain a clean, reproducible workflow.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.