Zero-Downtime Reloads in PM2 Cluster Mode: Configuration and Pitfalls
Configure PM2 cluster mode for true zero-downtime reloads: ecosystem file settings, graceful shutdown signals, ready protocol, and the pitfalls that cause silent failures.
18 Oct 2025, 11:47 UTC

The Problem and the Takeaway
When you deploy updates to a Node.js service running under PM2, the default pm2 restart kills all workers simultaneously. Active requests fail, TCP connections drop, and clients see 5xx errors. Use pm2 reload in cluster mode instead: it signals each worker to finish in-flight requests, then spawns a replacement before terminating the old process. The load balancer (PM2's built-in round-robin socket distribution) never loses the listening socket, so zero connections are dropped.
How Zero-Downtime Reload Works
PM2's cluster mode (-i max or a fixed instance count) forks a master process that holds the listening socket. Workers accept connections via round-robin. On pm2 reload <app>:
- The master sends
SIGINT(configurable viashutdown_signal) to one worker. - That worker stops accepting new connections, finishes its request queue, then exits.
- The master forks a fresh worker with the updated code.
- Repeat for each worker until all run the new version.
The socket stays open in the master the entire time. Clients never see a connection reset.
Worked Example: Ecosystem File for a Production Service
Create ecosystem.config.js in your project root. This single file makes reload behavior explicit and repeatable across environments.
module.exports = {
apps: [{
name: 'api',
script: 'dist/server.js',
instances: 'max', // one worker per CPU core
exec_mode: 'cluster',
wait_ready: true, // worker must call process.send('ready')
listen_timeout: 8000, // ms to wait for ready signal
shutdown_signal: 'SIGINT', // graceful signal (default)
kill_timeout: 5000, // force SIGKILL after this ms
max_memory_restart: '500M', // auto-restart if heap exceeds
env: {
NODE_ENV: 'production',
PORT: 3000
},
env_staging: {
NODE_ENV: 'staging',
PORT: 3001
},
log_date_format: 'YYYY-MM-DD HH:mm:ss Z',
merge_logs: true,
max_log_size: '10M', // rotate at 10 MB
retain_logs: 30 // keep 30 rotated files
}]
};
Key fields for zero-downtime reload:
wait_ready: true+listen_timeout: The master waits forprocess.send('ready')from the new worker before routing traffic. Add this in your bootstrap:if (process.send) { process.send('ready'); }kill_timeout: If a worker ignoresSIGINTpast this threshold, PM2 sendsSIGKILL. Set it longer than your longest request.instances: 'max': Uses all CPU cores. Replace with a number (e.g.,4) for container environments with CPU limits.
Deploying with Zero Downtime
Run on the target host (requires the same user that owns the PM2 daemon, typically your deploy user):
# Initial start
pm2 start ecosystem.config.js --env production
# Subsequent deploys
pm2 reload ecosystem.config.js --env production
# Verify all workers restarted
pm2 list
pm2 logs api --lines 50
Expected check: pm2 list shows status: online for all instances, restarts incremented by 1, and uptime reset to seconds. No errored or stopped entries.
Common Mistakes and Limits
1. Version-Specific Flag Changes
PM2 6.x renamed --update-env behavior and altered reload exit codes. Always verify flags for your version:
pm2 -v
pm2 reload --help | head -30
If you script deploys in CI, pin the PM2 version in package.json ("pm2": "^5.3.0") and test the reload flow in staging.
2. Ecosystem File Syntax Errors
A trailing comma or unquoted key in ecosystem.config.js causes silent fallback to defaults. Validate before deploy:
node -c ecosystem.config.js && echo 'syntax ok'
# or
pm2 start ecosystem.config.js --dry-run 2>&1 | head -20
Run this in your CI pipeline. A syntax error here prevents a broken production deploy.
3. Default Log Rotation May Fill Disk
The built-in rotation triggers at 10 MB (max_log_size) and keeps unlimited files unless you set retain_logs. For low-traffic apps, a single 10 MB file can take weeks to rotate, but a burst can generate gigabytes in hours. The example above sets retain_logs: 30 and explicit max_log_size. Adjust based on log volume:
# Check current log sizes
ls -lh ~/.pm2/logs/
# Rotate manually if needed
pm2 flush
pm2 reloadLogs
4. Worker Not Sending 'ready' Signal
If wait_ready: true but your app never calls process.send('ready'), the master times out after listen_timeout (default 3000 ms) and kills the worker. You'll see repeated spawn/kill cycles in pm2 logs. Fix: add the ready signal after your server listen() callback or database connection resolves.
5. Container CPU Limits vs. instances: 'max'
In Kubernetes or Docker with --cpus=2, 'max' reads host CPU count, not the cgroup limit. You'll spawn too many workers, causing CPU throttling. Explicitly set instances: 2 (or use a startup script that reads /sys/fs/cgroup/cpu.max).
Verification Checklist
- Version check:
pm2 -vmatches your tested version. - Cluster count:
pm2 listshows expected instance count after start. - Graceful reload: Run
pm2 reload apiwhile hitting the endpoint withhey -n 1000 -c 50 http://localhost:3000/health. Zero failed requests. - Log rotation: Generate >10 MB logs (
for i in {1..20000}; do curl -s http://localhost:3000/; done), confirmapi-out-1.logappears andapi-out.logshrinks. - Memory guard: Simulate leak (
setInterval(() => arr.push(new Array(1e6)), 100)), verify PM2 restarts the worker when heap crosses500M.
When Not to Use Cluster Mode Reload
- Stateful workers: If workers hold WebSocket connections, in-memory caches, or singleton DB pools that don't survive
SIGINT, reload drops that state. Use a shared Redis store or sticky sessions instead. - Single-instance apps:
exec_mode: 'fork'(default) has no master socket holder;reloadbehaves likerestartwith downtime. - Blue/green or canary needs: PM2 reload is all-or-nothing. For traffic splitting, use a reverse proxy (nginx, Traefik) with multiple PM2 apps on different ports.
Rollback
If a reload introduces a regression, pm2 reload ecosystem.config.js --env production with the previous commit's artifact restores the prior version using the same zero-downtime path. No separate rollback command needed.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.