Handling Slave Channels During Master Connection Failures
OpenSSH does not transparently re-establish a master connection to recover active slave channels. If the master connection times out or is severed, all multiplexed slave channels are terminated immediately. There is no native mechanism within OpenSSH to migrate an existing slave session's state or buffered data to a new master connection.
Likely Explanation of Behavior
Because ControlMaster relies on a single underlying TCP socket, the slave channels are logical abstractions (SSH channels) within that one encrypted tunnel. When the master connection fails—whether due to a network timeout or a ServerAliveCountMax expiration—the TCP socket is closed. Since the state of the slave channel (including any unacknowledged data in the local buffer) exists only within the context of that specific socket, the data is discarded when the connection terminates.
If a client script attempts to "retry" a command by launching a new SSH process while the master is reconnecting, it is initiating a new slave channel on a new master connection. It is not resuming the previous session. Consequently, any data that was buffered but not acknowledged by the server before the first master connection failed may be sent again by the application layer, potentially leading to duplicate writes.
Prevention and Mitigation
To prevent connection drops and the resulting data uncertainty, focus on maintaining the stability of the master connection rather than attempting to recover slave channels:
- Keep-Alives: Use
ServerAliveInterval and ServerAliveCountMax to ensure the master connection remains active during idle periods.
- Persistence: Use
ControlPersist to keep the master connection open in the background, reducing the frequency of new handshake events.
- Application-Level Idempotency: Since SSH cannot guarantee atomicity across connection failures, ensure that the commands being sent are idempotent (i.e., executing them twice has the same effect as once).
Verification Steps
You can verify this behavior by observing the socket lifecycle:
- Start a master session:
ssh -M -S /tmp/ssh-socket user@host
- Start a slave session:
ssh -S /tmp/ssh-socket user@host "sleep 100"
- Manually kill the master process or disrupt the network.
- Observe that the slave session terminates immediately with a "Connection closed" or "Broken pipe" error.
Diagnostic Detail Needed: Are you using a wrapper script or a specific automation tool (e.g., Ansible, Fabric) to handle the retries? The retry logic of the calling application determines whether duplicate writes occur, as OpenSSH itself does not perform transparent session migration.