Direct answer
1. Sticky sessions alone are not sufficient for cross-node rooms. Sticky sessions prevent Session ID unknown for the lifetime of a connection on a given node, but they do not make rooms work across nodes. Without an adapter, a server on node A only knows about sockets connected locally. Emitting to a room from node A will not reach clients that are connected to node B. If you can accept that during a rolling deploy some clients may miss broadcasts from nodes they are not pinned to, sticky sessions alone can be a temporary mitigation. Any requirement for reliable server-initiated broadcasts to all room members across the fleet effectively mandates a Redis adapter.
2. WebSocket-only reduces polling churn but does not remove the need for affinity. Engine.IO still performs an HTTP handshake before the upgrade. The handshake creates the session ID in memory on the node that handled it. With sticky sessions, the subsequent WebSocket frames stay on that node via TCP affinity. Without sticky sessions, a reconnect or the initial handshake can land on a different node and produce HTTP 400 Session ID unknown. Forcing transports:['websocket'] removes the long-polling fallback and the risk of a polling request bouncing mid-session, but it does not make the session portable across nodes.
3. Cross-node delivery can be verified without full production rollout. Start two server processes on different hosts/ports behind the same adapter configuration, connect clients to each node, join the same room, and emit from one node. With no adapter you will see the event only on the local node. With the Redis adapter installed on both processes you will see delivery to the remote client. Logging socket.id and io.of('/').adapter.rooms per process is a practical check.
Confirmed facts
- Engine.IO session state is in-memory per Node.js process. A packet referencing a session ID that is not present on the receiving process returns HTTP 400
Session ID unknown. - The
@socket.io/redis-adapter provides inter-node pub/sub for events and room membership. It does not store the Engine.IO handshake session; the handshake must still reach a node that holds the session. - Sticky session / session affinity keeps a client’s HTTP requests pinned to one backend. It prevents mid-session 400 errors as long as the backend stays alive.
Likely explanation for your deploy symptom
During a rolling restart the load balancer routes a polling request for an existing session to a new pod that never performed the handshake. The new pod has no in-memory session, so Engine.IO rejects it with Session ID unknown. This is consistent with non-sticky routing and with node termination while clients remain pinned to the old node.
Minimal steps for this case
- Enable cookie-based sticky sessions on the load balancer for the Socket.IO path, e.g.,
io cookie or JSESSIONID equivalent. Verify the affinity persists across the HTTP upgrade. - Keep default transports unless you have verified all clients can complete a WebSocket upgrade through your proxies. If you force websocket-only, test reconnect behavior behind restrictive proxies.
- If you need broadcasts to reach clients on any node, add the Redis adapter to all instances with the same Redis instance and namespace. Do not rely on the adapter to fix
Session ID unknown; keep sticky sessions for the handshake phase. - During deploys, drain connections on the old node before termination or use a short connection TTL to force clean reconnects to the new fleet.
One diagnostic detail that changes the recommendation: what Socket.IO major version are you running and what sticky mechanism does your load balancer provide, cookie-based affinity vs IP hash? Adapter initialization and sticky cookie names differ between v2 and v3/v4, and IP hash can break with NAT or mobile clients.