Confirm the failure
First check the cluster membership and service catalog to see if the node is marked as failed.
consul members
consul catalog nodes
If the node appears with status "failed" or is absent from the members list, the failure is recorded.
Review agent logs
Look for log entries indicating the node could not reach peers; the source notes that "Unreachable nodes will be marked as failed."
Determine next steps
Because Consul cannot distinguish a network partition from an agent crash, both situations are treated the same.
If the agent process is not running
- Verify the Consul process is stopped (e.g.,
ps -ef | consul or kill -KILL consul_pid).
- When the node is ready to return, start the agent again with your usual method, for example:
consul agent -config-file=/etc/consul.d/server.hcl
- The agent will join the cluster via gossip; existing servers will begin replicating to the new node if it is a server.
If the agent is still running but isolated
- Use the
force-leave command to mark the node as left:
consul force-leave <node-id>
- This puts the instance in a left state as long as the agent is not alive, allowing the cluster to remove it via the reaping process.
- After the node is removed, you can restart the agent and let it re‑join the cluster.
Verify recovery
After the agent has rejoined, check the logs for the "EventMemberJoin" message and confirm that the node appears in the catalog as active. The source states that "Existing servers begin replicating to the new node if the agent is a server," indicating successful recovery.
Note: If you cannot determine whether the agent process is running on the node, ask for access to the host (e.g., SSH) to check the process status before deciding between a restart or force‑leave.