timeout waiting for pod to be ready when restoring Kubeflow Pipelines workflow-controller pod
27.5K reputation · 16 Oct 2023, 19:40 UTC
The goal is to determine whether Velero‑based filesystem backups or CSI‑driven volume snapshots provide a reliable, application‑consistent restore path for Kubeflow Pipelines artifact stores (MLflow and MinIO) without triggering the 'timeout waiting for pod to be ready' error during workflow‑controller pod recovery.
The uncertainty stems from Velero restore failures that occur when the restored pod depends on a PVC that is not yet bound, because Velero does not wait for the underlying CSI snapshot to reach ReadyToUse before creating the pod. While CSI snapshots can reduce restore latency, they require identical storage class names between source and target clusters and careful handling of the KFP metadata database exclusion guidance, which varies across Kubeflow versions.
Which backup approach—Velero filesystem backup or CSI volume snapshot—ensures application‑consistent restores of KFP artifact stores without causing the workflow‑controller timeout? What storage class alignment steps are necessary when using CSI snapshots for the workflow‑controller PVC? How can post‑restore validation confirm that artifact links in MinIO and MLflow remain intact and accessible in the KFP UI?