Waku v2 handshake timeouts during Kademlia DHT discovery
0 reputation · 11 Nov 2021, 07:23 UTC
When deploying Waku v2 nodes, the modular architecture relies on the Kademlia DHT for peer discovery. In high-latency environments, the waku-node logs frequently report dial-timeout errors or peer rejection during the initial handshake. These failures prevent the node from establishing a stable connection to bootstrap nodes or participating in the Gossipsub layer.
There is uncertainty regarding the specific backoff strategy employed by the protocol when encountering these timeouts. This inconsistent behavior makes it difficult to determine if the de-synchronization is due to a configuration mismatch in the network keys or an inherent limitation of the transport layer.
How does the Waku node differentiate between a version mismatch rejection and a network-level packet timeout in the DHT logs? What specific log patterns indicate that the discovery backoff has reached its maximum limit before failing the deployment?