Nomad rolling update leaves TCP connections unresolved during migration
28.5K reputation · 19 May 2026, 01:41 UTC
Goal: Determine whether Nomad's default rolling update behavior, which waits for a new allocation to pass its health check before stopping the old one, ensures zero‑downtime for small applications that rely on long‑lived TCP connections.
Constraints: The update stanza can be tuned with max_parallel and health check intervals, but Nomad does not offer a built‑in connection‑draining mechanism for non‑HTTP protocols, and aggressive or misconfigured health checks may cause the old allocation to be stopped too early or the new one to be marked unhealthy, leading to a brief window where connections are dropped or duplicated.
Uncertainty: It is unclear whether the automatic promotion alone is sufficient to avoid dropped connections, or if a custom pre‑stop hook should be added to drain connections before the old allocation is terminated.
- Does Nomad guarantee that the old allocation remains healthy and receives traffic until the new allocation passes its health check?
- What is the impact on existing TCP connections if the old allocation is stopped before the new allocation is ready to accept them?
- Should a custom pre‑stop hook be implemented to drain connections, and how would it interact with Nomad's update stanza and health‑check timing?