'Failed to acquire leadership' in YB-Master logs: what retry behavior should operators expect?
0 reputation · 31 May 2021, 14:16 UTC
YugabyteDB's YB-Master relies on Raft-based leader election to choose the cluster leader. During startup or after a leader transition, a master that cannot win the election logs the documented entry Failed to acquire leadership, indicating the node could not take the cluster-leader role.
The best-documented cause is quorum loss: when fewer than a majority of master nodes are reachable because of crashes or a network partition, no candidate can win. Less clear is the process-level behavior after this entry. Whether yb-master exits, keeps retrying with backoff, or continues as a follower is not fully documented and appears to depend on version-specific configuration flags, with log wording also varying between 2.x and 3.x releases.
A further diagnostic ambiguity is that network latency or restrictive firewall rules can surface as election timeouts, making a slow or blocked network look like a genuine leadership failure.
- What determines whether a YB-Master retries after failing to acquire leadership, and which configuration flags control that policy?
- How does this behavior differ between YugabyteDB 2.x and 3.x?
- How can a quorum-loss failure be distinguished from a network-induced timeout using the master logs alone?