Moleculer Built-in Fault Tolerance: Policy Composition Without Library Sprawl
Moleculer ships timeout, retry, circuit‑breaker and bulkhead middlewares that compose declaratively via a policy object—eliminating the need for external resilience libraries. Understand the execution order, state sync, and cancellation wiring.
28 Apr 2026, 08:32 UTC

The problem: resilience without the baggage
You're running Node.js microservices. One calls a flaky third‑party API, another hits a database that occasionally stalls. The typical solution is to pull in opossum for circuit breaking, async-retry for retries, p-timeout for timeouts, and maybe bottleneck for concurrency limits. Then you wire them together, configure each differently per service, and hope the composition order doesn’t bite you at 3 AM.
How the policy stack composes
When you declare:
policy: { timeout: 3000, bulkhead: 10, retry: 3, circuitBreaker: { threshold: 0.5, windowTime: 60000 } }
The middleware chain executes in this order: timeout → bulkhead → retry → circuit‑breaker. That means:
- A request first hits the timeout wrapper. If the handler exceeds 3 seconds,
ERR_TIMEOUTis thrown (the original handler keeps running unless you hookctx.cancel()). - If the timeout passes, the bulkhead checks concurrent executions. At 10 concurrent calls, further requests get
ERR_BULKHEAD_REJECTEDimmediately—before any retry logic sees them. - Only then does retry kick in, with exponential backoff and optional jitter. You can supply a
checkfunction to retry only on specific error codes (e.g.,err.code === 'ETIMEDOUT'). - Finally, the circuit breaker records the outcome. When the failure rate crosses the threshold (0.5 = 50 %) over
windowTime, it opens and short‑circuits subsequent calls.
This order is intentional: you want to fail fast on timeout and overload before burning retry budget, and you want the circuit breaker to see the final outcome after retries.
Circuit breaker state: local by default, shared on demand
By default, each Moleculer node maintains its own circuit‑breaker state (closed/open/half‑open). In a cluster, that means Node A can hammer a failing dependency while Node B has already opened its circuit—until the transporter syncs the state.
With a shared transporter (NATS, Redis, MQTT) you can enable synchronization via the nodeID filter in the circuitBreaker options. The syncInterval (default 10 s) controls how often nodes exchange state. Tighten it for tighter coordination, but know that every sync adds broker chatter.
circuitBreaker: {
threshold: 0.5,
windowTime: 60000,
minRequestCount: 5,
syncInterval: 5000 // ms, requires shared transporter
}
Caution: threshold is a failure rate, not an absolute count. Low‑traffic actions may never trip or trip too eagerly. Set minRequestCount to require a minimum sample size before evaluating the rate.
Worked example: a flaky downstream HTTP service
Imagine a payments service that calls an external provider. The provider occasionally hangs (timeout), returns 5xx (retryable), or gets overwhelmed (bulkhead). Here’s a minimal action with all four policies:
const axios = require('axios');
module.exports = {
name: 'payments',
actions: {
charge: {
policy: {
timeout: 2000,
bulkhead: 5,
retry: {
count: 3,
delay: 200,
factor: 2,
check: (err) => err.code === 'ETIMEDOUT' || (err.response && err.response.status >= 500)
},
circuitBreaker: {
threshold: 0.5,
windowTime: 60000,
minRequestCount: 10
}
},
async handler(ctx) {
const controller = new AbortController();
ctx.on('cancel', () => controller.abort());
const resp = await axios.post('https://provider.example/charge', ctx.params, {
signal: controller.signal,
timeout: 1900 // slightly under policy timeout
});
return resp.data;
}
}
}
};
What happens under load?
- Concurrent calls > 5 → immediate
ERR_BULKHEAD_REJECTED(HTTP 503 from API gateway). - Provider hangs 2.1 s → timeout fires,
ERR_TIMEOUT, retry logic seesETIMEDOUTand retries up to 3 times with 200/400/800 ms delays. - Provider returns 503 → retry logic catches it (status ≥ 500) and retries.
- After 10+ requests with > 50 % failure rate → circuit opens, further calls fail fast with
ERR_CIRCUIT_OPENuntil half‑open probe succeeds.
The AbortController wiring is critical: without it, a policy timeout leaves the outbound HTTP request running, leaking sockets and connection‑pool slots. Moleculer’s timeout does not automatically cancel downstream I/O—you must integrate with the HTTP client’s cancellation mechanism.
Trade‑offs and limitations
- Bulkhead before retry: A transient spike that fills the bulkhead returns 503 immediately rather than queuing. If you prefer queueing, wrap the action with a small async queue (e.g.,
p-queue) inside the handler. - Timeout ≠ cancellation: As shown above, you must wire
ctx.cancel()toAbortControlleror axios cancel tokens. This is a common source of resource leaks. - Low‑traffic circuit breaker: With
minRequestCount: 10, an action called once per minute will never trip. Tune per action. - Hot‑reload caveat: Policy options merge at service load time. Changing policies on a running service requires a full broker restart or the
hotReloadplugin; otherwise the old middleware chain persists. - Sync interval latency: In multi‑node deployments, a failing dependency can be hammered independently by each node until
syncIntervalpropagates the open state. For tighter coordination, use a shared Redis store and lowersyncInterval(at the cost of more broker traffic).
Observability: metrics you can actually use
Each policy emits Moleculer events ($metrics.trace, $metrics.stat) that the official moleculer-prometheus module scrapes. Key counters:
moleculer_circuit_breaker_state_changes_total— tracks open/half‑open/closed transitionsmoleculer_bulkhead_rejections_total— counts bulkhead rejectionsmoleculer_retry_attempts_total— counts retry attempts
Add the module, expose /metrics, and you have dashboards and alerts without custom instrumentation.
What to try next
- Create a minimal service with the policy config above and an action that sleeps 200 ms then throws. Call it concurrently 10 times; watch broker logs for
ERR_TIMEOUT,ERR_BULKHEAD_REJECTED, retry attempts, and circuit‑open events. - Enable
moleculer-prometheusand verify the counters increment as expected. - If you run multiple broker instances, add a NATS transporter, trigger failures on one node, and confirm the circuit opens on the second node within your
syncInterval. - Wrap an outbound axios call with
AbortControllerlinked toctx.cancel(); verify the underlying TCP connection closes on timeout (inspect withnetstator Wireshark).
Moleculer’s built‑in policies won’t replace a dedicated service mesh for cross‑cutting concerns like mTLS or distributed tracing, but for action‑level resilience they eliminate a whole category of dependency sprawl. The declarative policy object keeps configuration next to the code it protects, and the fixed composition order means fewer surprises at 3 AM.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.