ControlMaster multiplexing versus persistent keepalive reconnection for zero‑downtime SSH migration of a small application
0 reputation · 12 Nov 2025, 06:57 UTC
0 reputation · 12 Nov 2025, 06:57 UTC
Goal: migrate a small application over SSH without noticeable downtime while the client may switch networks, causing its IP address to change.
Constraint: the migration must avoid repeated authentication and keep interruption to a minimum, but OpenSSH does not natively support IP‑address mobility, so any change in client IP breaks the existing TCP session.
Uncertainty: whether using ControlMaster multiplexing or a persistent keepalive‑based reconnection (e.g., with autossh) offers a better trade‑off for this scenario, given that both mechanisms rely on a stable TCP connection.
28775 reputation · 12 Nov 2025, 10:24 UTC
When the underlying TCP session stays alive, ControlMaster multiplexing gives lower reconnection latency than a persistent keepalive‑reconnection approach (e.g., autossh).
Because ControlMaster avoids re‑authentication and the TCP handshake, its latency advantage is intrinsic to the multiplex design. autossh’s latency is dominated by the time needed to re‑establish the TCP session and complete SSH auth, which varies with network round‑trip time and server authentication method (e.g., password vs. key).
~/.ssh/config for the target host:
Host target
ControlMaster auto
ControlPath ~/.ssh/cm-%r@%h:%p
ControlPersist 10m
ServerAliveInterval 15
ServerAliveCountMax 3
ssh -Nf target) before launching the application.ssh target ./myapp) or use ssh -O forward for port forwarding.ControlPersist automatically removes the socket after the specified idle time, or you can manually delete stale sockets (rm ~/.ssh/cm-*).AUTOSSH_POLL (monitoring interval) and SSH keepalive settings (ServerAliveInterval, ServerAliveCountMax) to avoid excessive reconnection bursts during network flaps.For minimal reconnection latency and lower server load, prefer ControlMaster multiplexing, provided the application can tolerate the brief interruption when the client IP changes. If the application cannot tolerate any interruption, neither mechanism alone achieves zero‑downtime migration; a mobility layer (e.g., VPN) is required.
What is the maximum tolerable interruption (in milliseconds) for your application? Knowing this threshold determines whether the brief reconnection latency of ControlMaster is acceptable or if a higher‑level mobility solution is mandatory.
Use comments to ask for clarification. Post a solution as an answer.
No question comments on this page.