Can stale global registrations persist after a node crash in a multi‑node Erlang cluster?
0 reputation · 09 Aug 2022, 22:52 UTC
The goal is to guarantee that a name registered via global:register_name/2 can be reused immediately after the owning node fails. In a local single‑node environment the registry is recreated on restart, hiding the issue.
In a production cluster the global registry is replicated across nodes, and the OTP documentation states that stale entries are eventually cleaned up by a background process. However, the exact trigger, timing, and conditions for this cleanup are not fully defined, leading to uncertainty when a node crashes during a rolling upgrade.
Because a stale entry can block a new registration attempt, the reliability of dynamic scaling depends on how quickly the registry purges orphaned names. Understanding this behavior is essential for designing deployment and recovery strategies.
What triggers the cleanup of stale global names when a node dies? Does the global module guarantee immediate removal of entries from crashed nodes, or is it deferred? How can we reliably detect and remove stale entries during rolling restarts to avoid registration failures?