Why GenServer and Supervision Trees Are the Practical Choice for Stateful Elixir Services
GenServer and supervision trees give fault‑isolated, stateful processes with automatic restart and state recovery—no external orchestrator needed. A worked example shows per‑device telemetry ingest with DynamicSupervisor.
14 Jan 2026, 10:13 UTC

The problem: stateful services that don't melt down under failure
You're building a service that holds per-client state—WebSocket connections, game sessions, device telemetry. When one client's process crashes, you don't want the whole node to restart. You also don't want to write custom restart logic, circuit breakers, or state-recovery boilerplate for every new domain object. Elixir's GenServer and supervision trees solve this by making failure isolation a first-class language feature, not a library add-on.
What GenServer actually gives you
A GenServer is a thin wrapper around OTP's gen_server behaviour. It turns a module into a long-running process with a mailbox, a state variable, and a fixed set of callbacks: init/1, handle_call/3 (synchronous), handle_cast/2 (asynchronous), handle_info/2 (plain messages), terminate/2, and code_change/3. The runtime guarantees sequential message processing, so you never race on internal state. Timeouts and caller monitoring are built in—no extra dependencies.
Supervision trees: 'let it crash' with boundaries
A Supervisor starts child processes and defines what happens when they exit. Strategies like :one_for_one (restart only the crashed child), :one_for_all (restart all siblings), and :rest_for_one (restart the crashed child and any started after it) let you model blast radius explicitly. The supervisor itself is supervised, forming a tree that reaches the application root. If a child exceeds max_restarts within max_seconds, the supervisor gives up and propagates the failure upward—preventing restart storms.
Worked example: per-connection GenServer under a DynamicSupervisor
Imagine a telemetry ingest service. Each device opens a TCP connection handled by its own GenServer. A DynamicSupervisor starts children keyed by device ID so you can look them up later.
# lib/telemetry/device_server.ex
module Telemetry.DeviceServer do
use GenServer
@impl true
def start_link(device_id, initial_state \\ %{}) do
GenServer.start_link(__MODULE__, {device_id, initial_state}, name: via_tuple(device_id))
end
defp via_tuple(device_id), do: {:via, Registry, {Telemetry.Registry, device_id}}
@impl true
def init({device_id, state}) do
# Recover persisted state if the process restarted
persisted = Telemetry.Store.load(device_id)
{:ok, Map.merge(state, persisted)}
end
@impl true
def handle_cast({:metrics, metrics}, state) do
new_state = Telemetry.Aggregator.merge(state, metrics)
# Persist asynchronously; don't block the mailbox
Task.start_link(fn -> Telemetry.Store.save(device_id, new_state) end)
{:noreply, new_state}
end
@impl true
def handle_call(:snapshot, _from, state), do: {:reply, state, state}
@impl true
def code_change(_old_vsn, state, _extra), do: {:ok, state}
end
# lib/telemetry/supervisor.ex
defmodule Telemetry.Supervisor do
use DynamicSupervisor
def start_link(opts \\ []) do
DynamicSupervisor.start_link(__MODULE__, opts, name: __MODULE__)
end
@impl true
def init(_opts) do
DynamicSupervisor.init(strategy: :one_for_one,
max_restarts: 5,
max_seconds: 10)
end
def start_device(device_id, initial_state \\ %{}) do
child_spec = %{
id: device_id,
start: {Telemetry.DeviceServer, :start_link, [device_id, initial_state]},
restart: :transient, # restart only on abnormal exit
shutdown: 5_000
}
DynamicSupervisor.start_child(__MODULE__, child_spec)
end
end
When a device's process crashes (say, a malformed payload triggers a FunctionClauseError in handle_cast), the DynamicSupervisor restarts only that child. The new init/1 call loads the last persisted snapshot from Telemetry.Store. Sibling device processes keep running uninterrupted. No external orchestrator, no custom health checks.
Trade-offs you'll hit in production
- Single-threaded mailbox: CPU-bound work in
handle_callorhandle_castblocks all messages for that device. Offload heavy computation to aTaskpool or aGenStageproducer. - Large state = slow restarts: If a device accumulates megabytes of history in the GenServer,
init/1replay becomes a latency spike. Move bulky data to ETS, Redis, or a time-series DB; keep only the hot aggregate in the process. - Restart intensity defaults bite: The default
max_restarts: 3, max_seconds: 5is too aggressive for flaky upstream dependencies. Tune per service criticality; a telemetry ingest might toleratemax_restarts: 10, max_seconds: 60. - No automatic replication: In a clustered deployment, GenServer state lives on one node. Network partitions mean that state is unavailable. For consensus-critical data, layer a Raft library (e.g.,
Ra) or CRDTs on top. - Synchronous call cascades: A chain of
GenServer.call/2with default 5-second timeouts can deadlock under load. Prefercast/2for fire-and-forget, or explicitGenServer.call(server, msg, 2000)with a short timeout and a{:reply, reply, new_state, timeout}return tuple.
Verifying the behaviour without a debugger
You can confirm the supervision contract in a live system using built-in tooling:
- Start
:observer.start()(or:observer_clion headless nodes) and inspect the Applications tab. You'll see the supervision tree, each child's restart count, and current memory. - Send a crash signal:
Process.exit(pid, :kill)from an IEx session attached to the node. Watch the supervisor restart the child and the restart counter increment. - Query state without side effects:
:sys.get_state(pid)returns the current term. Combine with:telemetryevents emitted fromhandle_castto graph mailbox depth and processing latency.
For release upgrades, build a new version with a changed callback, run mix release, and deploy with release_handler. The running GenServer processes will invoke code_change/3 automatically—test this in staging before relying on zero-downtime deploys.
Closing: start small, measure, then scale the pattern
GenServer plus supervision isn't a silver bullet—it's a disciplined way to isolate failure and recover state without leaving the language runtime. Start by modeling one domain entity (a session, a connection, a job) as a GenServer under a DynamicSupervisor. Add persistence in init/1. Observe restart behaviour under real load with :observer. Only then decide if you need clustering, Raft, or a different storage backend. The pattern scales because the primitives compose, not because they hide complexity.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.