Eliminating Deployment Downtime with PM2 Cluster Mode
Learn how to use PM2 Cluster Mode to eliminate deployment downtime. This guide covers ecosystem configuration, rolling restarts, and the necessity of stateless architecture.
15 Mar 2026, 16:47 UTC

The Deployment Gap Problem
When you restart a Node.js application using a standard process manager or a simple npm start, there is a window of time where the process is shutting down and the new one is booting up. During this gap, any incoming HTTP requests are rejected, resulting in 502 Bad Gateway or Connection Refused errors for your users.
The solution is to move from a single-process model to a cluster model. By using PM2's Cluster Mode, you can maintain a pool of workers behind a built-in load balancer, allowing you to replace workers one-by-one without ever stopping the service entirely.
How Cluster Mode Enables Zero-Downtime
In Cluster Mode, PM2 acts as a master process. It doesn't run your application code itself; instead, it spawns multiple worker processes (instances) of your app. The master process handles the incoming network connections and distributes them across the workers using a round-robin strategy.
When you trigger a reload, PM2 performs a rolling restart. It sends a signal to the first worker to shut down, waits for it to exit, and immediately starts a new worker with the updated code. Only after the new worker is online does PM2 move to the next instance. Because other workers remain active throughout the process, the application continues to serve traffic.
Implementing a Cluster Configuration
While you can start a cluster from the command line, using an ecosystem.config.js file is the professional standard for ensuring consistency across environments.
Step 1: Create the Ecosystem File
Run this command in your project root to create the configuration file:
touch ecosystem.config.js
Step 2: Define the Cluster Settings
Add the following configuration. This example assumes a Node.js app starting from server.js:
module.exports = {
apps : [{
name: "api-service",
script: "./server.js",
// 'max' utilizes all available CPU cores
instances: "max",
exec_mode: "cluster",
// Ensures the process has time to finish requests before being killed
kill_timeout: 3000,
env: {
NODE_ENV: "production",
PORT: 3000
}
}]
};
Step 3: Launch and Reload
Run the following commands on your server with the user permissions required to manage the Node.js process:
# Initial start
pm2 start ecosystem.config.js
# To update code with zero downtime
pm2 reload api-service
Verification and Diagnostics
To verify the cluster is active, run pm2 list. You should see multiple rows for api-service, each with a unique ID but the same app name, all showing a status of online.
To test the zero-downtime capability, you can use a load-testing tool like hey or ab (Apache Benchmark) to send a steady stream of requests while issuing the reload command:
# Run this in a separate terminal to monitor requests
hey -z 30s http://localhost:3000/health
While that is running, execute pm2 reload api-service. If configured correctly, the load tester should report zero failed requests, though you may see a slight increase in latency for the specific requests handled by the worker being cycled.
Critical Limitations and Trade-offs
Cluster mode is not a silver bullet. It introduces two primary architectural constraints:
- Statelessness: Since requests are distributed across different processes, you cannot store session data or cached variables in local memory. If Worker A stores a user session in a local variable, Worker B will not have access to it. You must use an external store like Redis or a database for session management.
- The Kill Timeout: If your application has long-running requests (e.g., large file uploads), the default
kill_timeoutmight be too short. If PM2 force-kills a worker before it finishes a request, that specific user will experience a connection drop. Always align yourkill_timeoutwith your longest expected request duration.
Rollback Procedure
If a reload introduces a bug, you can revert to the previous stable version by redeploying the previous code commit and running the reload command again. Because PM2 manages the process state, if a new worker fails to start (crashes on boot), PM2 will stop the rollout, leaving the remaining old workers active to prevent a total site outage.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.