Migrating to Kubeflow Pipelines v2: Decoupling Compilation from Execution
Stop treating ML pipelines as opaque scripts. Learn how Kubeflow Pipelines v2 uses a static compilation model to turn Python DSL into portable Argo YAML for better CI/CD and Kubernetes native control.
13 Feb 2026, 18:21 UTC

The Problem: The "Black Box" Pipeline
In Kubeflow Pipelines (KFP) v1, the pipeline definition was often tightly coupled to the running environment. Creating a pipeline frequently felt like writing a script that had to be executed against a live API to see if it actually worked. This made CI/CD difficult: you couldn't easily version-control a "compiled" pipeline or validate the workflow structure without a functioning Kubeflow control plane.
The takeaway is that KFP v2 shifts the paradigm from dynamic script execution to a static compilation model. By compiling pipelines into portable YAML files (based on Argo Workflows), teams can treat ML pipelines as infrastructure-as-code, enabling GitOps workflows and predictable deployments.
The Shift to Component-Based Authoring
The most immediate change in v2 is the move away from ContainerOp toward the @dsl.component decorator. In v1, components were often opaque wrappers around container images. In v2, the SDK uses Python type annotations to define the contract between steps.
When you use @dsl.component, the KFP compiler extracts the function signature and requirements to create a standalone component specification. This removes the need for Python pickling—a common source of versioning headaches in v1—and replaces it with explicit Input and Output types. These components are portable; you can share a component YAML across different pipelines without redistributing the original Python source code.
Execution Model: Native Kubernetes Semantics
KFP v2 pipelines compile directly to Argo Workflow YAML. This means each step in your pipeline is a native Kubernetes Pod. This is a critical engineering win because it grants you direct access to Kubernetes primitives without needing an abstraction layer in the SDK.
- Resource Management: You can set CPU/Memory limits and requests directly on the component.
- Scheduling: Use node selectors, tolerations, and affinity to ensure heavy GPU training jobs land on specific hardware nodes.
- Retries: Pod-level retries are handled by the Argo backend, making the pipeline more resilient to transient infrastructure failures.
Parameter vs. Artifact Passing
A subtle but important optimization in v2 is how data moves between steps. Small parameters (typically under 10MB) are now passed via a lightweight JSON mechanism using /tmp/inputs/ and /tmp/outputs/. This reduces the latency and dependency on the MinIO/S3 artifact store for simple scalar values, though large datasets and models still require an external object store.
Worked Example: Defining a v2 Component
To use KFP v2, you need the kfp SDK version 2.x. The following example demonstrates a component that takes a string and outputs a processed value, compiled to a YAML file for GitOps deployment.
import kfp.dsl as dsl
from kfp import compiler
# Define a reusable component
@dsl.component(base_image="python:3.9")
def preprocess_data(input_text: str) -> str:
# Logic happens inside the container at runtime
return f"Processed: {input_text.upper()}"
# Define the pipeline workflow
@dsl.pipeline(name="text-processing-pipeline")
def my_pipeline(text_input: str = "hello world"):
# The output of the component is passed as a parameter to the next
proc_task = preprocess_data(input_text=text_input)
print(f"Pipeline output: {proc_task.output}")
# Compile the pipeline to a YAML file
if __name__ == "__main__":
compiler.Compiler().compile(my_pipeline, "pipeline.yaml")
Verification and Deployment
To verify the compilation, run the script locally. You will see a pipeline.yaml file generated in your directory. You can then deploy this using the KFP UI or via the CLI if you have argo installed:
# Run this on a cluster with Argo Workflow CRDs installed
# Required permissions: Cluster Admin or namespace-specific create permissions for workflows
argo submit pipeline.yaml -n kubeflow --watch
Risk: Note that compiler.Compiler().compile() only validates the syntax and structure. It does not check if the base_image exists in your registry or if the Kubernetes cluster has enough resources to schedule the pods. These errors will only appear at runtime.
Trade-offs and Limitations
The transition to v2 is not a drop-in replacement. Because the DSL surface area has changed—removing dsl.ContainerOp and dsl.PipelineParam—existing v1 pipelines must be rewritten. While a compatibility package exists, it is a shim and not a long-term solution for complex graphs.
Additionally, the @dsl.component decorator executes the function at compile time to extract the signature. If you place network calls or heavy computations inside the function body (outside of the logic intended for the container), they will run on your local machine during compilation, potentially leading to non-reproducible builds or failed CI pipelines.
Actionable Closing
If you are starting a new project or your current v1 pipelines have become a maintenance burden, migrate to KFP v2. To start, verify your backend version: kubectl get deployment ml-pipeline -n kubeflow -o jsonpath='{.spec.template.spec.containers[0].image}'. Ensure you are running a version compatible with KFP 1.8+ to support the v2 backend. Begin by converting your most stable utility functions into @dsl.component specs to build a library of reusable, versioned YAML components.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.