Zero-downtime migration of a joblib-persisted scikit-learn pipeline across library versions
0 reputation · 13 Jul 2026, 06:17 UTC
Scenario
A small Flask service loads a scikit-learn Pipeline (StandardScaler + LogisticRegression) from a joblib file at startup and serves predictions in production. The plan is to move from scikit-learn 1.3 to 1.5 without any request downtime, since the service has no maintenance window.
Constraints
The scikit-learn documentation states that models persisted with one version are not guaranteed to load or behave identically in another, because internal attributes and parameter defaults can change. There is no built-in model registry or versioning mechanism, and only a subset of estimators supports ONNX export via skl2onnx as a version-agnostic format. The current pipeline contains one custom transformer, which may not be exportable.
A dual-run approach (old and new model side by side behind a flag) is feasible, but it is unclear how to detect silent prediction drift between the two runtimes rather than only catching load-time AttributeError failures.
Questions
- Is loading a 1.3-era joblib artifact into a 1.5 runtime considered safe enough for a canary rollout, or should retraining on 1.5 be treated as mandatory?
- What is a practical way to compare predictions between the old and new model versions on live traffic before cutting over?
- Given a custom transformer in the pipeline, is ONNX export a realistic compatibility bridge, or does it introduce more divergence risk than it removes?