Gardener v1.87+ shoot API server certificate rotation: transient x509 failures when control plane restarts during renewal
0 reputation · 23 Jul 2024, 09:37 UTC
On Gardener v1.87+ shoots, certificate rotation via cert-manager updates the serving certificate secret, and the kube-apiserver is expected to pick it up through its TLS file watcher (Kubernetes >= v1.22) without a restart. The new CA bundle is then propagated to clients through the shoot's status credentials.
My concern is the ordering gap: if the kube-apiserver pod restarts (crash, node drain, or reconciliation-triggered rollout) after the secret is updated but before the file watcher reloads it, the API server may serve a certificate that clients no longer trust, producing intermittent x509: certificate signed by unknown authority errors. There appears to be no documented synchronization barrier ensuring all API server replicas have reloaded before the new CA bundle is published.
Assume a shoot on Kubernetes 1.28, Gardener v1.87, with manual rotation triggered via the gardener.cloud/operation=rotate-certificates annotation.
- Is the rotation-to-reload window formally bounded anywhere, or is it purely eventually consistent by design?
- Does Gardener sequence CA bundle publication after confirming API server reload, and if not, is this a known gap?
- Does concurrent kube-apiserver autoscaling during rotation widen this race?