Using Kubeflow Pipelines with Argo Workflows: Reproducible ML on Kubernetes
Learn how Kubeflow Pipelines uses Argo Workflows to run each component in its own pod, enabling fine‑grained resource control, versioning, and full reproducibility. A step‑by‑step example shows how to write, submit, and verify a minimal pipeline on a production‑ready cluster.
03 Jan 2026, 00:32 UTC

Problem: Reproducible ML Workflows on Kubernetes
Data scientists want to run experiments in a cloud‑native way, but they often struggle with “it works on my laptop, not on the cluster” problems. The root cause is usually a lack of isolation, version control, and a clear audit trail. Operations teams need a single source of truth for every run, and they want to enforce resource limits to avoid starving other workloads.
Thesis: Kubeflow Pipelines + Argo Workflows = Reproducible, Isolated, Scalable ML
Kubeflow Pipelines (KFP) turns each step of a machine‑learning workflow into an Argo Workflow step that runs in its own Kubernetes pod. The Python SDK generates the Argo YAML automatically, and the KFP UI stores run metadata in PostgreSQL. Together, they provide:
- Fine‑grained resource isolation via per‑step pod limits.
- Automatic versioning of pipeline runs and component images.
- Horizontal scaling by leveraging Kubernetes’ scheduler.
- Built‑in integration with popular ML frameworks.
How KFP Maps to Argo Workflows
When you submit a pipeline, the SDK translates each component into an Argo Container step. Each step runs in a separate pod that lives in the kubeflow namespace (or a custom namespace you specify). This isolation lets you:
- Set
cpuandmemorylimits per step. - Attach sidecar containers for logging or monitoring.
- Run steps in parallel or serially based on DAG dependencies.
Example of the generated Argo YAML for a two‑step pipeline:
apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata:
generateName: my-pipeline-
namespace: kubeflow
spec:
entrypoint: my-pipeline
templates:
- name: my-pipeline
dag:
tasks:
- name: preprocess
template: preprocess
- name: train
template: train
dependencies: [preprocess]
- name: preprocess
container:
image: myrepo/preprocess:1.0
command: ["python", "preprocess.py"]
- name: train
container:
image: myrepo/train:2.1
command: ["python", "train.py"]
Versioning and Reproducibility
Each pipeline run gets a unique run_id that the KFP UI uses to fetch logs, parameters, and results from PostgreSQL. Because component images are pinned to a specific tag, you can re‑run the exact same version by selecting the run in the UI or via the REST API:
- Open the run details page.
- Click Re‑run and choose the same image tags.
- The new run will have a fresh
run_idbut identical component images.
To verify reproducibility, run the pipeline twice with different image tags and confirm that the metadata reflects the tag differences:
# Submit run 1
kfp run --pipeline my_pipeline.yaml --image-tag 1.0
# Submit run 2
kfp run --pipeline my_pipeline.yaml --image-tag 1.1
# Check PostgreSQL metadata (example query)
SELECT run_id, component_name, image_tag FROM pipeline_runs WHERE pipeline_name='my_pipeline';
Scaling and Resource Isolation
Because each step is a pod, Kubernetes can schedule them across the cluster. You can set resource limits per step to avoid contention:
container:
image: myrepo/train:2.1
resources:
limits:
cpu: "4"
memory: "8Gi"
requests:
cpu: "2"
memory: "4Gi"
When multiple pipeline runs are active, the scheduler will spread them out, but you must still monitor cluster utilization. The KFP UI displays pod status and logs, and you can use kubectl top pod -n kubeflow to spot resource hotspots.
Concrete Example: Minimal Pipeline
Below is a step‑by‑step guide to write, submit, and verify a minimal pipeline that prints the Python version.
- Write the pipeline in Python:
- Deploy KFP (if not already deployed): Follow the official quick‑start guide to install Kubeflow Pipelines 1.8 on your cluster. Verify the Argo UI is reachable at
/argoin thekubeflownamespace. - Submit the pipeline via the REST API:
- Verify pod creation:
- Check logs:
- Inspect metadata in PostgreSQL: Run a query against the
runstable to confirm the run ID and step logs are stored.
from kfp import components, dsl
@components.func_to_container_op
def python_version():
import sys
print(sys.version)
@dsl.pipeline(name='Python Version Demo')
def demo_pipeline():
python_version()
if __name__ == '__main__':
from kfp.compiler import Compiler
Compiler().compile(demo_pipeline, 'demo_pipeline.yaml')
# Replace placeholders
PIPELINE_FILE=demo_pipeline.yaml
PIPELINE_NAME=python-version-demo
# Upload the pipeline
curl -X POST -H "Content-Type: application/json" \
-d "{\"name\":\"$PIPELINE_NAME\",\"pipeline_spec\":{\"pipeline_spec\":{\"pipeline_id\":\"$(uuidgen)\"}}}" \
http://localhost:3000/api/v1/pipelines
# Create a run (simplified; use actual API endpoints for production)
curl -X POST -H "Content-Type: application/json" \
-d "{\"pipeline_id\":\"$(cat $PIPELINE_FILE | base64)\",\"run_name\":\"$PIPELINE_NAME-run\"}" \
http://localhost:3000/api/v1/runs
kubectl get pods -n kubeflow -l "app.kubernetes.io/component=pipeline-run"
kubectl logs -n kubeflow <pod-name>
Trade‑Offs and Limitations
- RBAC Complexity: The pipeline needs permissions to create pods, services, and config maps in the target namespace. Mis‑configured RBAC can block execution.
- Resource Limits: Under‑provisioning leads to pod evictions; over‑provisioning wastes cluster capacity. Use
kubectl top nodeto monitor usage. - Version Compatibility: Newer KFP releases may rely on Argo features that older Argo installations lack. Keep both components on compatible versions.
- PostgreSQL Single Point of Failure: The metadata store is a single node by default. Scale it with replication or use a managed PostgreSQL service for production.
Actionable Closing
Adopting Kubeflow Pipelines with Argo Workflows gives you:
- Clear, auditable traces of every training run.
- Fine‑grained control over resources at the step level.
- Built‑in reproducibility via image pinning and run IDs.
- Scalability that leverages Kubernetes’ scheduler.
Next steps for your team:
- Audit your cluster’s RBAC to ensure pipeline pods can be created.
- Set up PostgreSQL replication for metadata resilience.
- Create a shared library of commonly used components (e.g., data‑loading, training) and pin them to stable tags.
- Automate pipeline submission via CI/CD to enforce versioning.
With these practices, you’ll turn ad‑hoc experiments into repeatable, production‑ready workflows that both data scientists and operations can trust.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.