Question
What timeout should the Couchbase upgrade service use before marking a node as upgrade‑failed?
Tasadduq BurneyownerOwner · Founder
29K reputation · 05 Jul 2022, 07:06 UTC
28K views0
When a Couchbase node upgrade fails, the upgrade service marks the node as ‘upgrade‑failed’ and leaves the cluster in a rebalance‑required state, but it does not automatically revert the node to the previous version. Administrators must decide how long to wait for the service to finish its internal retry before taking manual recovery actions, yet the documentation does not specify a timeout for this decision point. This uncertainty can lead to inconsistent interventions across clusters and versions, increasing the risk of premature downgrade or prolonged imbalance. What is the maximum safe interval administrators should wait before initiating a manual downgrade? How can log entries be used to determine whether the upgrade service is still retrying? Should a version‑specific timeout policy be adopted to ensure consistent recovery behavior?