Does Traefik Mesh support automatic rollback of TrafficSplit on health‑check failure?
0 reputation · 11 Jun 2023, 05:31 UTC
Traefik Mesh’s TrafficSplit resource enables weighted traffic shifting between service versions, but the current implementation does not automatically modify weights in response to health‑check failures or other observability signals. Operators must manually adjust the TrafficSplit or delete it to revert traffic, which can delay reaction to a failing version.
The unresolved design decision is whether the mesh should incorporate a built‑in rollback mode that watches endpoint health (e.g., via Kubernetes readiness probes or external metrics) and reduces the weight of unhealthy backends without operator intervention. This raises questions about the desired latency of rollback, the health‑check source, and how to avoid oscillations.
Does Traefik Mesh currently provide a mechanism to automatically adjust TrafficSplit weights based on health‑check status? If not, what are the recommended patterns for implementing external auto‑rollback? Are there any open issues or proposals that address automatic rollback for TrafficSplit?