Answer the question first
When a uWSGI upgrade fails, the safest recovery path is usually an immediate restart (SIGTERM then start the last known‑good binary) unless you can confirm that the failure occurred *before* any new workers were spawned and that the master process detected the error and rolled back via --reload-on-failure. In that narrow case a graceful reload (SIGHUP) may be sufficient.
Confirmed facts
- uWSGI master reacts to SIGHUP by finishing existing requests, then spawning new workers with the new binary/config while keeping the old workers alive until they finish.
- SIGTERM (or SIGQUIT) stops the master immediately; all workers are killed, allowing a clean start of a known‑good version.
- The
--reload-on-failure flag tells the master to revert to the previous binary/config if the reload handshake fails *before* the new workers accept connections.
- This rollback mechanism was added in uWSGI 2.0.1 (released early 2013) and has been stable since; older releases (<2.0) lack reliable rollback and may leave a partially upgraded state.
- If the upgrade changes the socket path (or socket type), the old workers keep listening on the old socket after a graceful reload, producing an orphaned socket that persists until those workers exit. A full restart is required to avoid the race condition.
Likely explanation
The master can only roll back when it detects a failure *during* the reload handshake: binary load error, config parse error, or worker init failure that causes the new workers to exit before they call accept(). If the failure happens after workers have started accepting requests (e.g., a runtime error in your application), the master treats the reload as successful and will not revert.
Steps for this case
- Check the uWSGI log for the reload attempt. Look for messages like "Unable to load plugin", "Invalid configuration", or worker exit codes before the "spawned" line.
- If the log shows a pre‑accept failure and you have
--reload-on-failure enabled, you can safely send SIGHUP; the master will have already rolled back.
- If the log shows workers started and then crashed, or if you cannot confirm the failure point, send SIGTERM to the master, verify all workers have exited (
ps -fu www-data | grep uwsgi), then start the last known‑good binary.
- After the action, verify recovery via the stats endpoint (
curl http://localhost:9191/stats) or uwsgitop: worker count should match the expected number, error rates (5xx) should drop to baseline, and response times should stabilize within a few seconds.
- If the upgrade changed the socket path, always perform the immediate restart (SIGTERM) path; graceful reload will leave the old socket bound until the old workers finish.
One missing diagnostic detail that would change the recommendation
Confirm whether the reload failure occurred *before* any new workers began accepting connections (i.e., during binary load/config parsing). If you can verify this, a graceful reload with --reload-on-failure may be sufficient; otherwise, proceed with an immediate restart.