Scaling Node.js Apps with PM2 Cluster Mode: When and How to Use It
Learn how PM2’s cluster mode spreads Node.js workloads across CPU cores, verifies load‑balancing, and handles zero‑downtime reloads—plus the memory and state limits you must watch.
29 Aug 2026, 07:37 UTC

Problem: Your Node.js API hits a CPU ceiling
You have a simple HTTP service built with Node.js that works fine under low traffic, but as request volume grows the single‑threaded event loop becomes a bottleneck. Adding more CPU cores to the host does not help unless the application can spread work across them.
Thesis: PM2’s cluster mode gives you automatic, zero‑code‑change horizontal scaling across CPU cores, with graceful reloads, as long as your service is stateless.
How cluster mode works
When you start an app with PM2 in cluster mode, a master process is spawned. The master forks a number of worker processes equal to the instances setting (or max to match the detected CPU cores). Incoming TCP connections are handed to workers using a round‑robin scheduler, so each worker gets a fair share of requests. Each worker runs its own V8 instance, meaning memory usage scales roughly linearly with the worker count.
Worked example: enabling cluster mode and verifying load‑balancing
Create a minimal Express server that logs its process ID on each request:
// server.js const express = require('express'); const app = express(); app.get('/', (req, res) => { res.send(`Hello from worker ${process.pid}`); }); const PORT = process.env.PORT || 3000; app.listen(PORT, () => console.log(`Listening on ${PORT}`));Start the app with PM2, asking for four instances (adjust to your core count):
# Run in your project directory pm2 start server.js -i 4You need permission to execute
pm2(typically just your user account). No sudo is required unless you are binding to a privileged port (<1024).Verify that requests are distributed:
# From another terminal, send 12 rapid requests for i in {1..12}; do curl -s http://localhost:3000; echo; doneYou should see the worker IDs rotate, e.g.:
Hello from worker 12345 Hello from worker 12346 Hello from worker 12347 Hello from worker 12348 Hello from worker 12345 ...
If the same ID appears repeatedly, check that
pm2 listshows the expected number of workers and that no external load balancer is interfering.Test zero‑downtime reload:
# Open a long‑running request (e.g., streaming) curl -s http://localhost:3000/stream & # Reload the app without dropping connections pm2 reload server.jsThe streaming request should finish normally while you see both old and new worker PIDs appear briefly in
pm2 list.
Trade‑offs and limitations
- Memory usage: Each worker duplicates the application code and its own heap. On a container with limited RAM, setting
-i maxcan cause OOM kills. Monitor memory withpm2 monitorpm2 listand compare the RSS column before and after scaling. - State affinity: In‑memory session stores, caches, or WebSocket connections are not shared between workers. If your app relies on such state, externalize it (Redis, a database, or a dedicated message broker) before using cluster mode.
- Sticky sessions: PM2’s built‑in round‑robin does not provide IP‑based affinity. If you need sticky sessions, place an external load balancer (NGINX, HAProxy) in front of PM2 or use the
pm2-proxymodule.
Actionable closing
If your Node.js service is stateless and you have spare CPU cores, enable PM2 cluster mode with a simple -i flag or an instances field in ecosystem.json. Verify load‑balancing by checking rotating worker IDs in logs or curl responses, and watch memory growth with pm2 monit. When you need to update code, use pm2 reload for zero‑downtime rollouts. Keep an eye on RAM usage and externalize any in‑memory state to reap the full benefits of horizontal scaling.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.