How can stateful connections be managed during a Cloud Run zero-downtime migration?
0 reputation · 11 Feb 2020, 21:30 UTC
0 reputation · 11 Feb 2020, 21:30 UTC
The goal is to perform a zero‑downtime migration of a small application by deploying a new Cloud Run revision and gradually shifting traffic via traffic splitting, keeping the service URL stable.
Because Cloud Run does not provide session affinity and treats each request independently, long‑lived connections (e.g., WebSocket, persistent HTTP, or database pools) that are bound to a specific revision may be terminated when traffic is moved, and any in‑memory session state is lost unless externalized.
What strategies exist for externalizing session state so that both revisions can access it? Can connection‑draining or keep‑alive mechanisms be used to allow existing connections to finish before they are terminated? Is it acceptable to rely on client‑side reconnect logic, or must the application be redesigned to be connection‑stateless?
29275 reputation · 12 Feb 2020, 04:06 UTC
Use Cloud Run’s traffic‑splitting feature to shift a small percentage of requests to the new revision while keeping the old revision alive, allowing existing stateful connections to finish, and externalize any session state so both revisions can access it.
Cloud Run routes traffic to multiple revisions simultaneously; as long as the container instance that holds a connection stays alive, the connection survives. A gradual traffic shift lets old connections drain naturally while new requests go to the new revision.
Cloud Run’s traffic splitting does not terminate existing instances when the split changes; it only stops sending new traffic to them, so in‑flight connections can continue until they close or idle‑timeout.
What is the observed connection idle timeout (or keep‑alive setting) for your stateful service? If the timeout is shorter than the expected migration window, you may need to adjust it or rely more heavily on externalized state and client‑side reconnect logic.
Use comments to ask for clarification. Post a solution as an answer.
29,275 reputation · 12 Feb 2020, 02:20 UTC
One premise worth correcting: Cloud Run does offer optional session affinity (gcloud run deploy --session-affinity), but it's best-effort and pins requests to a container instance, not a revision. It breaks whenever that instance scales down or is reclaimed, so it may reduce reconnect churn in steady state but is no migration guarantee.
The practical consequence: client-side reconnect logic is mandatory regardless. Cloud Run's per-service request timeout applies to WebSocket connections too, so connections are recycled on a fixed clock even with no deployment — five minutes by default, configurable higher. Check your service's actual value with gcloud run services describe. If reconnect-with-resumption against the externalized session store is already routine, a traffic shift becomes just another reconnect event rather than a special case.
For the drain side: instances receive SIGTERM before shutdown, so the old revision's shutdown handler should stop accepting work, send WebSocket close frames, and close database pools. The grace window before SIGKILL is short — verify the current documented value rather than hard-coding assumptions.