Immediate SIGKILL after kill_timeout vs Graceful SIGTERM Handling: Choosing the Right Nomad Shutdown Strategy
28.5K reputation · 25 Apr 2020, 03:07 UTC
When stopping a Nomad job allocation, the scheduler first sends SIGTERM and then waits the duration defined by kill_timeout before escalating to SIGKILL if the task is still running. For rolling updates, the stop_timeout in the update stanza determines how long Nomad waits for a canary allocation to stop before proceeding.
The goal is to choose a shutdown approach that balances update speed with the risk of leaving stale tasks or corrupting application state, while considering that Nomad does not automatically forward SIGTERM to descendant processes.
Uncertainty remains about whether to rely solely on Nomad’s built‑in timeout values or to introduce a wrapper/sidecar that propagates the termination signal to child processes, and how the two timeout settings interact during a canary update.
What kill_timeout value provides enough time for typical workloads to flush buffers without delaying updates excessively? Should a sidecar be employed to forward SIGTERM to child processes, or can the task itself handle the signal adequately? How does stop_timeout interact with kill_timeout when a canary fails to stop within the allotted interval?