Implementing Artifact Caching in Kubeflow Pipelines to Reduce Training Costs
Learn how to implement and verify artifact caching in Kubeflow Pipelines to eliminate redundant computations and reduce cloud infrastructure costs.
07 Oct 2026, 15:48 UTC

The Cost of Redundant ML Computations
In machine learning workflows, data preprocessing and feature engineering steps are often computationally expensive and time-consuming. Without a caching mechanism, every execution of a Kubeflow Pipeline (KFP) re-runs every component, even if the input data and parameters remain unchanged. This leads to wasted GPU/CPU hours and slower iteration cycles.
The solution is Artifact Caching. By enabling caching, KFP checks the metadata store to see if a component with the exact same inputs and code version has already executed. If a match is found, KFP skips the execution and injects the previous output directly into the next step.
Prerequisites
- A running Kubeflow cluster (v2.x recommended) with the KFP SDK installed in your local environment.
- A configured Persistent Volume Claim (PVC) or Object Store (like S3 or GCS) linked to the ML Metadata (MLMD) store to persist artifacts.
- Cluster-admin permissions or a ServiceAccount with permissions to create
PipelineRunandTaskCustom Resources.
Configuring Caching in the Pipeline DSL
Caching is managed at the component or pipeline level using the KFP Python SDK. You can enable caching globally for a pipeline or selectively disable it for steps that must run every time (e.g., a step that fetches the latest live data from an API).
import kfp
from kfp import dsl
@dsl.component
def preprocess_data(input_path: str, output_dataset: dsl.Output[dsl.Dataset]):
# Expensive data processing logic here
with open(output_dataset.path, 'w') as f:
f.write("Processed data content")
@dsl.component
def train_model(dataset: dsl.Input[dsl.Dataset]):
# Training logic using the processed dataset
print(f"Training on {dataset.path}")
@dsl.pipeline(name="Caching Example")
def ml_pipeline():
# Step 1: Preprocessing (Cache enabled by default in KFP 2.x)
preprocess_task = preprocess_data(input_path="s3://bucket/raw_data.csv")
# Step 2: Training (Explicitly disable cache for this step to ensure fresh training)
train_task = train_model(dataset=preprocess_task.outputs['output_dataset'])
train_task.set_caching_options(False)
Deployment and Execution
Run the following commands in your local terminal where the SDK is configured to point to your Kubeflow cluster:
- Compile the pipeline:
Check: Verify thatkfp.compiler.Compiler().compile(ml_pipeline, 'pipeline.yaml')pipeline.yamlcontains theapiVersionandkind: Pipelinedefinitions. - Upload and Run: Use the Kubeflow UI to upload
pipeline.yamland trigger a run.
Verifying Cache Hits and Misses
To confirm that caching is functioning, perform a comparative test between two consecutive runs:
| Action | Run 1 (Baseline) | Run 2 (Identical Inputs) |
|---|---|---|
| Preprocess Step | Executed (Full Runtime) | Cached (Near-instant) |
| Train Step | Executed (Full Runtime) | Executed (Fresh Run) |
| UI Indicator | Green checkmark | Grey/Blue "Cached" icon |
If the preprocess_data step executes again in Run 2 despite no changes, check your storage class. If the PVC used for the ML Metadata store is ephemeral or misconfigured, the system cannot retrieve the artifact hash, forcing a re-computation.
Diagnostic Decision: When to Disable Caching
Caching is not always desirable. Use the following logic to decide when to call .set_caching_options(False):
- Dynamic Inputs: If a component reads from a database or API where the data changes but the
input_pathstring remains the same. - Non-Deterministic Logic: If the component uses random sampling without a fixed seed.
- External State: If the component relies on an external environment variable or system state not tracked by KFP.
Rollback and State Reset
Since caching relies on the ML Metadata (MLMD) store, you cannot "undo" a cache for a single run without changing the input parameters. To force a complete refresh of all cached components, you must either:
- Change a parameter value (e.g., change
input_pathfromdata_v1todata_v1_refresh). - Manually clear the MLMD database (High risk: this deletes lineage for all pipelines in the cluster).
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.