Diagnosing and Fixing Erlang Process Mailbox Overflow
A diagnostic guide for identifying and resolving Erlang process mailbox overflows, covering root cause analysis, recon-based checks, and back-pressure strategies.
05 Oct 2025, 16:54 UTC

The Problem: Unbounded Mailbox Growth
In the BEAM VM, process mailboxes are unbounded by default. When a producer sends messages faster than a consumer can process them, the message_queue_len grows monotonically. This leads to a cascade of failures: increased memory pressure, higher garbage collection (GC) latency, and eventually, a node crash due to Out-Of-Memory (OOM) errors.
The primary takeaway is that mailbox growth is a symptom of a throughput mismatch or a logic block. Solving it requires identifying whether the process is too slow, stuck, or ignoring specific messages.
Diagnostic Matrix: Root Causes
| Symptom | Likely Cause | Key Metric to Check |
|---|---|---|
| Queue grows steadily; CPU is high | Producer exceeds consumer throughput | message_queue_len vs reductions |
| Queue grows; CPU is low/idle | Selective receive failure (unmatched messages) | current_function and code clauses |
| Queue grows; Process is 'busy' | Blocking synchronous call or NIF | current_stacktrace / Dirty schedulers |
| High memory; Low queue count | Mailbox fragmentation / Large binaries | message_queue_data (off-heap) |
Step-by-Step Runtime Investigation
Perform these checks in order using the Erlang shell or a remote shell (ssh) to minimize impact on the production node. This guide assumes the use of the recon library, the industry standard for BEAM diagnostics.
1. Identify the Offenders
Start by finding which processes are hoarding messages. Run this in the shell:
% List top 10 processes by message queue length
recon:proc_window(message_queue_len, 10, 1000).
If a few PIDs show consistently increasing lengths, they are your primary targets.
2. Inspect Process State
Once you have a PID, examine its current status and memory footprint:
% Replace Pid with the actual process identifier
erlang:process_info(Pid, [message_queue_len, memory, status, current_function]).
If status is running but message_queue_len is climbing, the process is working but too slowly. If status is sleeping while the queue grows, the process is likely ignoring the messages being sent to it.
3. Trace Message Handling
If the process is running but the queue persists, check if it is getting stuck in a specific function or failing to match a receive clause:
% Trace the handle_info or handle_cast calls for a GenServer
recon:trace(module_name, handle_info, 10).
Look for a lack of "catch-all" clauses. If a process uses selective receive (e.g., receive {ok, Msg} -> ... end) without a fallback, unmatched messages stay in the mailbox forever.
Targeted Fixes
Scenario A: Throughput Mismatch
If the producer is simply too fast, you must implement back-pressure. The BEAM does not provide this natively for gen_server:cast.
- Switch to synchronous calls: Replace
castwithcallto force the producer to wait for a response. - Use GenStage: Implement a demand-driven pipeline where producers only send data when consumers request it.
- TCP-style Windowing: Implement a credit-based system where the producer stops sending after X messages until an ACK is received.
Scenario B: Blocking Operations
If the process is stuck in a Native Implemented Function (NIF) or a heavy CPU task, it blocks the mailbox processing.
- Dirty Schedulers: Move heavy CPU work to dirty schedulers using
erlang:spawn_opt/4with[dirty_cpu]or mark NIFs asenif_schedule_dirty_nif. - Offloading: Move blocking I/O to a
poolboyworker pool or aTask.Supervisor.
Scenario C: Selective Receive Leaks
If the process is ignoring messages, ensure every receive block or gen_server callback has a default handler:
% Fix: Add a catch-all to prevent mailbox accumulation
handle_info(_UnexpectedMsg, State) ->
{noreply, State}.
Escalation and Safety Thresholds
Engage on-call engineering or initiate a node drain if any of the following thresholds are met:
- Queue Depth: Any single process queue exceeds 100,000 messages for more than 30 seconds.
- Memory Pressure: Total VM memory exceeds 80% of system RAM, with >30% residing in process heaps.
- Scheduler Saturation: Scheduler utilization > 95% for 5+ minutes with a run-queue > 2x the number of available schedulers.
Verification and Rollback
Verification: After deploying a fix, monitor the queue length over a 10-minute window. The message_queue_len should either stabilize or trend downward as the process clears the backlog.
Rollback: If back-pressure implementation causes producer timeouts or deadlocks, revert the change to the previous stable version. Note that reverting the code will not clear the existing accumulated mailbox; you may need to restart the affected processes to purge the queue.
Cautionary Notes
- Crash Dumps: Triggering a crash dump via
kill -USR1can create files several gigabytes in size. Ensure sufficient disk space before analysis. - Off-Heap Binaries: Large binaries referenced in the mailbox prevent GC reclamation. Clearing the mailbox is the only way to free this memory.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.