PM2 Cluster Mode or Fork Mode? A Decision Guide for Node.js Deployments
When one pm2 process is enough, what cluster mode demands from your app's state handling, and how to configure and verify zero-downtime reloads.
20 Sept 2025, 01:52 UTC

You are about to put a Node.js service into production with pm2, and one setting decides how your hardware gets used: exec_mode. Fork mode (the default) runs exactly one process, so your JavaScript executes on a single core no matter how many the machine has. Cluster mode starts one worker per core and spreads HTTP requests across them. The short version: fork mode for scripts, background workers, and low-traffic internal apps; cluster mode for production HTTP servers — but only if you are willing to move in-memory state such as sessions, caches, and WebSocket connections out of the process.
Assumptions: pm2 5.x with Node.js 18 or later on Linux. Run the commands on the production host, in your project root, as the same OS user that owns the pm2 daemon (pm2 talks to that user's ~/.pm2 directory; no root required). The behavior described here is long-standing, but confirm it on your versions with the checks at the end.
The decision and its constraints
Three questions settle it:
- Is the app a network server? Cluster mode is built on Node's
clustermodule, which shares one listening socket between worker processes. It only makes sense for apps that callserver.listen(). Batch jobs and queue consumers belong in fork mode. - Do you keep state in process memory? Workers are separate processes with separate heaps. A login session held in a local variable exists in one worker only. If you cannot move that state to Redis or a database, stay in fork mode, or pin users to workers with sticky sessions at your reverse proxy.
- How much memory is free? Each worker is a full Node process. If your app idles at 200 MB, four workers cost roughly 800 MB before any traffic arrives.
Fork vs cluster at a glance
| Aspect | Fork mode (default) | Cluster mode |
|---|---|---|
| Processes | One | One per core with instances: 'max', or any fixed count |
| Request routing | None needed — single process | PM2 round-robins HTTP connections across workers |
| Zero-downtime reload | No — restart closes the process and its socket | Yes — workers restart one at a time |
| Memory | One heap | One heap per worker; total use scales with count |
| In-memory state | Works, but a crash loses it | Not shared between workers; needs Redis, a database, or sticky routing |
| Good fit | Workers, scripts, small internal APIs | Production HTTP servers and APIs |
What cluster mode buys you, and what it costs
Gains. Throughput scales with cores because pm2 hands each incoming connection to the next worker in rotation (round-robin). Deployments stop costing you availability: pm2 reload replaces workers one at a time while the listening socket stays open, so clients do not see connection refused errors. A memory leak in one worker takes down only that worker — pm2 restarts it, or you can cap it with max_memory_restart — while the others keep serving.
Costs. Shared state must move to an external store such as Redis. Logs interleave across workers, so tag log lines with an identifier — pm2 exposes each worker's index as the NODE_APP_INSTANCE environment variable (0, 1, 2, …). Long-lived WebSocket connections are the rough edge: round-robin suits short HTTP requests, and a socket pinned to one worker breaks when that worker restarts. For WebSockets, terminate them at a reverse proxy with sticky routing (for example nginx with ip_hash) or keep the socket layer in fork mode. Debugging also gets harder: you are reasoning about N processes instead of one.
Configuring cluster mode in ecosystem.config.js
Define the mode declaratively rather than with command-line flags. Create ecosystem.config.js in your project root:
module.exports = {
apps: [{
name: 'web-api',
script: './src/index.js',
exec_mode: 'cluster', // 'fork' is the default
instances: 'max', // one worker per CPU core, or a fixed number such as 2
max_memory_restart: '500M', // restart a worker whose heap exceeds this
kill_timeout: 5000, // ms of grace for in-flight requests on shutdown
env: {
NODE_ENV: 'production'
}
}]
};
Start it from the project root on the server:
pm2 start ecosystem.config.js
Expected result: pm2 list shows one row per worker with the mode column reading cluster. If your app is slow to initialize (database pools, cache warm-up), add wait_ready: true and listen_timeout: 10000, and call process.send('ready') once the app can actually serve — otherwise pm2 may mark a worker online before it is.
Validating before you rely on it
- Instance count. Run
pm2 list. The number ofweb-apirows should match yourinstancesvalue — with'max', your CPU core count. - Work distribution. Run
pm2 monit, send some traffic, and confirm CPU time appears across several worker PIDs rather than piling onto one. - Zero-downtime reload. In a second terminal, stream requests at a cheap endpoint (replace the port and path with your app's):
Then runwhile true; do curl -s -o /dev/null -w '%{http_code}\n' http://127.0.0.1:3000/health sleep 0.5 donepm2 reload web-apiin the first terminal. You should see an unbroken stream of 200 responses. Connection refused or bursts of 5xx mean workers are dying before in-flight work finishes — checkpm2 logs web-api, raisekill_timeout, and make sure the app closes its server gracefully onSIGINT.
Limitations
- Cluster mode adds concurrency, not speed per request: a CPU-bound computation still blocks whichever worker receives it.
- Round-robin provides no session affinity; that responsibility stays with your reverse proxy.
- Total memory use scales with instance count — size the host before enabling
'max'.
Rolling back to fork mode
A reload cannot switch the execution mode. Edit the config (exec_mode: 'fork', instances: 1) and run pm2 delete web-api && pm2 start ecosystem.config.js. That removes all workers at once, so schedule it outside traffic peaks.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.