Do Couchbase alert thresholds survive node removal without a full rebalance?
0 reputation · 30 Mar 2021, 12:10 UTC
Goal
I want to configure threshold-based alerts in Couchbase Server (memory usage, disk usage, cache miss ratio) via Settings -> Alerts, route only Critical severity to PagerDuty, and throttle everything else so the team is not flooded with notifications.
Constraint
Our cluster topology changes fairly often, and nodes are sometimes removed under time pressure. I have seen indications that alert definitions stored through the UI or the /pools/default/alerts REST endpoint live in the cluster configuration, and that a forced node removal without a proper rebalance might drop them. If that is true, our carefully tuned thresholds and per-channel routing could silently disappear exactly during an incident.
I am also unsure how the global throttling defaults interact with per-alert overrides, and whether a maintenance window suppresses only notifications or also the alert state itself.
Questions
- Are alert definitions guaranteed to persist after a node is failed over and removed, or only after a clean rebalance?
- Does a per-alert throttle override fully replace the global throttle, or are both applied?
- Is exporting alert definitions via the REST API and re-applying them after topology changes the recommended safeguard?