Speeding Up Iterative ML Workflows with Kubeflow Pipelines v2 Step Caching
Learn how Kubeflow Pipelines v2 step caching skips unchanged ML workflow steps, saving time and cost, with a concrete three‑step example and guidance on when to disable caching.
28 Jun 2026, 23:01 UTC

When you tweak a hyperparameter or adjust a model architecture, the rest of your pipeline—data loading, preprocessing, feature engineering—often stays the same. Re‑running those unchanged steps wastes cluster time and increases cost. Kubeflow Pipelines (KFP) v2 addresses this with automatic step caching: if a component’s inputs, code, and container image haven’t changed, the platform reuses the previous run’s outputs instead of re‑executing the step.
How step caching works in KFP v2
Each pipeline step is a @dsl.component function that builds a container image and declares its inputs and outputs. When the KFP backend prepares a run, it computes a hash from:
- The component’s source code (including any imported modules).
- The exact container image digest used for the step.
- The runtime values of all input parameters and artifacts.
If that hash matches a previous successful execution, the step is marked as “cached” and its outputs are restored from ML Metadata. The step appears as skipped in the UI, and no new pod is scheduled.
Worked example: a three‑step ML pipeline
Consider a pipeline where only the training hyperparameters change between experiments. The first two steps—loading raw data and preprocessing—can safely be cached, while the training step must run fresh each time.
# pipeline_components.py
import kfp.dsl as dsl
import kfp
@dsl.component
def load_data(output_path: dsl.Output[dsl.Dataset]):
# Simulated: download or read a static dataset
with open(output_path.path, 'w') as f:
f.write('sample data')
@dsl.component
def preprocess(data: dsl.Input[dsl.Dataset], processed_path: dsl.Output[dsl.Dataset]):
# Simulated: a deterministic transformation
with open(data.path, 'r') as infile, open(processed_path.path, 'w') as outfile:
outfile.write(infile.read().upper())
@dsl.component
def train(processed: dsl.Input[dsl.Dataset], learning_rate: float, model_path: dsl.Output[dsl.Model]):
# Non‑deterministic: uses random seed, so we disable caching
import tensorflow as tf
import numpy as np
# Dummy model
model = tf.keras.Sequential([tf.keras.layers.Dense(1, input_shape=(1,))])
model.compile(optimizer=tf.keras.optimizers.SGD(learning_rate=learning_rate),
loss='mse')
x = np.array([[1.0], [2.0], [3.0]])
y = x * learning_rate # synthetic target
model.fit(x, y, epochs=2, verbose=0)
model.save(model_path.path)
@dsl.pipeline(name='hyperparameter-tuning', description='Cache load & preprocess, re‑run train')
def hp_pipeline(learning_rate: float = 0.01):
data_task = load_data()
prep_task = preprocess(data=data_task.output)
# Disable caching for the training step because it depends on randomness
train_task = train(processed=prep_task.output, learning_rate=learning_rate)
train_task.set_caching_options(False) # explicit opt‑out
if __name__ == '__main__':
kfp.compiler.Compiler().compile(hp_pipeline, 'hp_pipeline.yaml')
Where to run the snippet: on a local workstation or CI builder that has the KFP SDK installed and network access to your target Kubernetes cluster. You need read/write permissions to the cluster’s API server to create PipelineRuns.
Compiling and submitting the pipeline
# Install the SDK (choose a version that matches your cluster)
pip install "kfp==2.4.0" # adjust version as needed
# Compile the pipeline to the intermediate representation
python pipeline_components.py # produces hp_pipeline.yaml
# Submit a run (replace placeholders)
kfp.Client(host="https:///pipeline")\
.create_run_from_pipeline_package(
"hp_pipeline.yaml",
arguments={"learning_rate": 0.02},
experiment_name="hyperparameter-tuning",
namespace="")
Required permissions: the service account used by the KFP client must be able to create PipelineRun and TaskRun custom resources in the chosen namespace. If you lack these rights, ask your cluster administrator to grant the edit role on the namespace.
Expected check: after the first run finishes, open the KFP UI, locate the run, and observe that the load_data and preprocess steps show a “cached” badge (or are greyed out as skipped). The train step will display a normal execution status.
Trade‑offs and limitations
- Cache invalidation sensitivity: Changing the component’s source code, its base image, or any input parameter will produce a new hash and force a re‑run. This is desirable for correctness but can surprise developers who edit a helper function used by a cached step.
- External mutable data: If a step reads from a database, API, or shared filesystem that can change between runs, caching may return stale results. The safest approach is to disable caching for such steps with
set_caching_options(False), as shown for the training step. - Version drift: The caching semantics have evolved between KFP v1.x and v2.x, and even across v2 minor releases. Always verify the installed
kfppackage version against the documentation for your Kubeflow distribution. - Operational overhead: Using only the pipelines component (stand‑alone KFP) is lightweight, but the full Kubeflow platform adds Istio, cert-manager, and additional controllers. For teams that only need caching‑enabled pipelines, a standalone install reduces operational complexity.
Practical way to verify caching behavior
- Run the pipeline twice with identical arguments.
- In the UI, compare the
Started attimestamps: the first run will show timestamps for all three steps; the second run should show timestamps only for thetrainstep, while the other two display a dash or “cached” label. - To test invalidation, modify the
preprocesscomponent (e.g., add a logging line), bump its version, re‑compile, and submit a new run. You should see the preprocessing step execute again.
Actionable closing
If your team iterates on model hyperparameters or feature engineering, enable step caching in Kubeflow Pipelines v2 to cut unnecessary compute. Start by:
- Confirming the SDK version matches your cluster’s KFP backend.
- Writing deterministic components for data ingest and transformation.
- Explicitly disabling caching for any step that relies on randomness or external mutable state.
- Validating cache hits via the UI before expanding to larger pipelines.
By treating caching as a first‑class design decision, you reduce waste, speed up feedback loops, and keep costs predictable while preserving reproducibility through ML Metadata lineage.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.