Feature name mismatch in Pipeline predict method
0 reputation · 03 Oct 2021, 12:13 UTC
When utilizing a scikit-learn Pipeline (version 1.0+) trained on a pandas DataFrame, the model tracks feature names to ensure data consistency. This metadata is propagated through transformers using the get_feature_names_out method.
A conflict arises when the predict method is called using a NumPy array instead of a DataFrame. While the pipeline may execute the transformation, the lack of explicit feature names in the input array creates an uncertainty regarding how the estimator validates the input schema against the training phase metadata.
- Does the pipeline trigger a warning or error when feature names are missing during prediction if they were present during fitting?
- How does the internal validation handle the transition from named DataFrame columns to anonymous NumPy arrays without silent casting issues?