Nomad rolling update leaves TCP connections unresolved during migration
0 reputation · 19 May 2026, 01:41 UTC
0 reputation · 19 May 2026, 01:41 UTC
Goal: Determine whether Nomad's default rolling update behavior, which waits for a new allocation to pass its health check before stopping the old one, ensures zero‑downtime for small applications that rely on long‑lived TCP connections.
Constraints: The update stanza can be tuned with max_parallel and health check intervals, but Nomad does not offer a built‑in connection‑draining mechanism for non‑HTTP protocols, and aggressive or misconfigured health checks may cause the old allocation to be stopped too early or the new one to be marked unhealthy, leading to a brief window where connections are dropped or duplicated.
Uncertainty: It is unclear whether the automatic promotion alone is sufficient to avoid dropped connections, or if a custom pre‑stop hook should be added to drain connections before the old allocation is terminated.
Nomad’s default rolling update does not guarantee that long‑lived TCP connections stay alive during the switch; the old allocation can be stopped as soon as the new allocation passes its health check, and Nomad does not automatically drain connections for host‑network tasks.
max_parallel and the health‑check interval; Nomad waits for the new allocation to be marked healthy before stopping the old one.nomad node drain command performs an explicit drain step before stopping allocations.When the new allocation passes its health check, Nomad promotes it and immediately tells the client to stop the old task. If the old task is still accepting or holding TCP connections, those connections are reset because the process receives a SIGTERM and exits, unless the application itself handles graceful shutdown.
ss -tnp for ESTABLISHED states).max_parallel = 1 and keep the health check interval short enough to detect readiness but long enough to avoid flapping.ss -tnp | grep <port> and confirm ESTABLISHED connections drop to zero before the task stops.Are you using Consul Connect bridge networking for this job? The answer changes whether a custom prestop hook is required.
Use comments to ask for clarification. Post a solution as an answer.
No question comments on this page.