Diagnosing and Resolving Memory Bloat in Elixir Applications
A diagnostic guide to Elixir memory bloat: distinguish process growth from legitimately large data, find the hog processes, and apply targeted fixes for state, binaries, ETS, and recursion.
29 Jul 2025, 04:18 UTC

In Elixir, memory issues are rarely caused by traditional "leaks" where pointers are lost. Because the Erlang Virtual Machine (BEAM) manages memory per-process, memory bloat usually occurs when processes accumulate state or stay alive while holding large data structures. If a GenServer's state grows without being pruned, that memory is not returned to the operating system until the process terminates.
To fix the problem, first distinguish between high memory usage (data that is reachable and legitimately large) and process growth (processes or data that should have been released). Optimizing the wrong one wastes effort.
Identifying the Pattern
Before changing code, characterize the growth. Use this table to map observed symptoms to likely root causes:
| Symptom | Likely Cause | Diagnostic Metric |
|---|---|---|
| Process count increases and never drops | Zombie consumers or supervisors restarting crashed children | :erlang.system_info(:process_count) |
| Memory stays high after load decreases | Long-lived GenServer state or orphaned ETS tables | :erlang.process_info(pid, :memory) |
| Sudden spikes during specific tasks | Large lists or binary references held in state | :erlang.process_info(pid, :binary) |
| Memory grows linearly inside a loop | Non-tail recursion or unbounded accumulation | Observer Processes tab |
Step 1: Isolate the Memory Hogs
Start by identifying which processes consume the heap. In a development or staging environment, use the built-in observer GUI to visualize the process tree. Run this in an IEx session connected to the node (requires no special permissions on your own node, but on a remote node you need distributed Erlang access, which is a security-sensitive capability):
:observer.start()In the observer, open the Processes tab and sort by the Memory column. Look for processes consuming far more than your application baseline. On a remote server without a GUI, list the top memory consumers from a remote shell instead:
# Top 10 processes by memory (run in a remote iex shell)
Process.list()
|> Enum.map(fn pid ->
{pid, :erlang.process_info(pid, :memory)}
end)
|> Enum.sort_by(fn {_pid, {:memory, mem}} -> mem end, :desc)
|> Enum.take(10)Expected check: the output should show a small number of processes dominating memory. If memory is spread evenly across thousands of processes, you have a process-count problem, not a per-process state problem.
Step 2: Analyze Process State and Data
Once you have a suspicious PID, determine what it holds. For a GenServer, inspect its state — but be cautious in production, since requesting a large state can block the process and copy a lot of data to your shell:
:sys.get_state(pid)Check for these common Elixir-specific pitfalls:
- Large binaries: Binaries over 64 bytes live on a shared binary heap, referenced by processes. If a long-lived process holds even a small sub-binary reference to a large binary, the entire original binary cannot be garbage collected. Copy the slice with
:binary.copy/1to release the original. - Accumulated lists or maps: Appending to a list in GenServer state without truncation grows it indefinitely.
- Orphaned ETS tables: ETS tables are not tied to process garbage collection. If the owning process dies without deleting the table, the data remains in memory until the table is explicitly deleted or a new owner takes over. List tables with
:ets.all()and check sizes with:ets.info(table, :memory).
Step 3: Apply Fixes Tied to Findings
Prune long-lived state
If a GenServer tracks a cache or history, bound it with a sliding window or a time-to-live (TTL):
def handle_cast({:add, item}, state) do
# Keep only the most recent 1000 items
new_state = Enum.take([item | state], 1000)
{:noreply, new_state}
endEnsure tail recursion
If memory grows inside a recursive loop, confirm the recursive call is the last expression. Non-tail recursion keeps a stack frame per iteration:
# Not tail-recursive: the + happens after the call
def sum([h | t]), do: h + sum(t)
# Tail-recursive with an accumulator
def sum(list), do: sum(list, 0)
defp sum([], acc), do: acc
defp sum([h | t], acc), do: sum(t, acc + h)Clean up ETS and hibernate idle processes
Delete ETS tables in a supervisor's or owner's terminate/2 callback, or transfer ownership deliberately. For GenServers that idle between bursts holding large state, return {:noreply, state, :hibernate} so the BEAM garbage-collects and shrinks the heap while the process waits. Hibernate trades a small wake-up cost for memory reclamation; avoid it for processes handling high-frequency messages.
Limitations and Verification
Manual garbage collection via :erlang.garbage_collect(pid) can confirm whether memory is reclaimable, but calling it routinely hurts throughput — treat it as a diagnostic, not a fix.
To verify a fix, record :erlang.memory(:total) before and after a representative load test. Healthy behavior: total memory rises under load and returns near baseline after load stops and GC runs. If it ratchets upward across repeated cycles, a leak remains. Also compare :erlang.system_info(:process_count) before and after; a stable count confirms processes are terminating as expected.
Escalation Criteria
Escalate beyond application-level fixes when: memory grows even with a stable process count and bounded states (suspect a NIF or port driver leaking outside the BEAM heap — check :erlang.memory(:system) versus :processes); fragmentation is suspected (inspect allocator stats with :erlang.system_info(:allocator)); or growth only appears under production traffic you cannot reproduce locally, in which case capture a crash dump or use recon's :recon_alloc tooling on a live node with appropriate access.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.