Tracking Kubeflow TF Training Runs with MLMD: A Working Setup and Its Pitfalls
Wire a Kubeflow Pipelines TensorFlow training component into MLMD so runs, parameters, and model artifacts are tracked durably — and avoid the ephemeral-storage mistake that silently erases metadata.
12 Apr 2026, 11:09 UTC

The short answer
Kubeflow Pipelines can wrap a TensorFlow training job as a pipeline component and record each run's inputs, outputs, and artifacts in MLMD (ML Metadata), the metadata store that powers lineage in the Kubeflow UI. The setup works reliably if you do three things: declare the training step as a container component, persist the MLMD database on a mounted volume, and pass data paths through environment variables rather than literals. The most common failure is a run that trains fine but leaves no metadata, almost always because the MLMD database lives on ephemeral pod storage.
This guide assumes Kubeflow Pipelines 2.x on a Kubernetes cluster you administer, with kubectl access and a default StorageClass that can provision PersistentVolumeClaims (PVCs). Exact component names and UI paths vary by distribution, so verify against your installed version.
How the pieces fit
MLMD is a metadata store (backed by SQLite or MySQL) that records executions, artifacts, and the links between them. When a pipeline runs, each step writes entries describing what it consumed and produced. The Kubeflow UI reads this store to render run details and lineage graphs. If the store's file or connection is lost between runs, you lose the history even though the training itself succeeded.
The data flow looks like this:
- You submit a compiled pipeline (YAML or via the KFP SDK).
- The pipeline orchestrator schedules the TF training pod and opens a connection to the MLMD backend.
- The training container runs your script; the orchestrator records the execution and its artifacts in MLMD.
- The UI queries MLMD to show the run, its parameters, and output artifacts.
A minimal worked configuration
The example below is a self-contained pipeline definition using the KFP v2 container-component style. It runs a TensorFlow training script and relies on the pipeline backend's MLMD connection for metadata. Adjust image tags and paths to your environment; this is a template, not a tested artifact.
# pipeline.yaml (illustrative KFP v2 component spec)\ncomponents:\n comp-tf-train:\n executorLabel: exec-tf-train\n inputDefinitions:\n parameters:\n data_path:\n parameterType: STRING\n epochs:\n parameterType: NUMBER_INTEGER\n outputDefinitions:\n artifacts:\n model:\n artifactType:\n schemaTitle: system.Model\n schemaVersion: 0.0.1\ndeploymentSpec:\n executors:\n exec-tf-train:\n container:\n image: tensorflow/tensorflow:2.15.0\n command: [python, /opt/train.py]\n args:\n - --data\n - {inputParameter: data_path}\n - --epochs\n - {inputParameter: epochs}\n - --model-out\n - {outputPath: model}\n env:\n - name: DATA_ROOT\n valueFrom:\n configMapKeyRef:\n name: training-config\n key: data_root\n resources:\n cpuLimit: \"4\"\n memoryLimit: 16G\n cpuRequest: \"2\"\n memoryRequest: 8G\npipelineInfo:\n name: tf-training-with-metadata\nroot:\n dag:\n tasks:\n tf-train:\n taskInfo:\n name: tf-train\n componentRef:\n componentName: comp-tf-train\n inputs:\n parameters:\n data_path:\n runtimeValue:\n constant: /data/train\n epochs:\n runtimeValue:\n constant: 10\nschemaVersion: 2.1.0\nsdkVersion: kfp-2.xKey decisions in this spec:
- ConfigMap for paths.
DATA_ROOTcomes from a ConfigMap so the same pipeline compiles and runs identically across clusters. Hard-coded paths are the top cause of \"works on my cluster\" failures. - Resource requests and limits. TF training pods are frequent victims of namespace quotas. Without explicit requests, the scheduler may place the pod on a node that cannot sustain it, or the pod may be evicted mid-run.
- Typed output artifact. Declaring
modelas asystem.Modelartifact is what lets MLMD record it with a meaningful type, which the UI renders as a first-class artifact rather than an opaque path.
Persisting MLMD so metadata survives
In a default Kubeflow deployment, MLMD is served by the metadata-grpc deployment backed by MySQL, or by SQLite on a PVC in lighter installs. The failure mode to guard against is the SQLite file (or MySQL data directory) sitting on container-local storage: when the pod restarts, all run history vanishes.
Check where your metadata store keeps its data:
# Run on your workstation with cluster-admin or kubeflow namespace read access\nkubectl get pods -n kubeflow -l component=metadata-grpc\nkubectl describe pod -n kubeflow -l component=metadata-grpc | grep -A5 VolumesLook for a volume backed by a persistentVolumeClaim, not emptyDir. If you see emptyDir, patch the deployment or Helm values to attach a PVC before running anything you care about. Risk: editing the metadata deployment restarts it, briefly interrupting the Pipelines API; do it outside active runs.
Limits and common mistakes
- Shared SQLite across concurrent writers. SQLite-backed MLMD does not tolerate concurrent writes well. If you run multiple pipeline instances against the same SQLite file, metadata can corrupt or silently drop. Use the MySQL-backed deployment for anything beyond a single-user test cluster.
- Scale ceilings are version-dependent. Published numbers for maximum runs per experiment or recommended database sizes vary by Kubeflow release and backend. Treat any specific figure you read (including older blog posts) as unverified until you check the docs for your installed version or observe behavior in your own cluster under load.
- Artifacts without declared types. If your component writes files without declaring them as typed outputs, training still succeeds but lineage in the UI is incomplete. Declare every output you want tracked.
- Assuming metadata equals reproducibility. MLMD records what ran; it does not pin your data. If
/data/trainchanges between runs, two runs with identical metadata produced different models. Version your dataset path (e.g., include a date or hash) to make runs actually comparable.
Verifying the setup end to end
- Submit the pipeline from the Kubeflow UI or with
kfp client create_run_from_pipeline_package. - Confirm the training pod reaches
Completed:kubectl get pods -n kubeflow-user-example-com(namespace name depends on your profile). - Open Pipelines → Runs in the UI and select your run. You should see the execution, its input parameters, and the
modelartifact with typesystem.Model. - Restart the metadata pod (
kubectl rollout restart deployment/metadata-grpc-deployment -n kubeflow, adjusting the deployment name to your install) and refresh the UI. If the run history is still there, your metadata is on persistent storage. If it disappears, fix the volume before proceeding.
That last restart test is the cheapest way to catch the ephemeral-storage mistake before it costs you real experiment history.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.