Implementing Fault-Tolerant Worker Patterns with Erlang OTP Supervisors
Learn how to implement Erlang OTP supervisors to ensure fault tolerance. This guide covers restart strategies, child specifications, and how to prevent infinite restart loops in concurrent systems.
30 Jun 2026, 17:35 UTC

The Problem: Preventing Cascading Failures in Concurrent Systems
In a distributed or highly concurrent system, a single unhandled exception in a worker process can lead to silent failures or orphaned resources. Without a management layer, you must manually track process IDs (PIDs) and implement complex retry logic within your business code, which mixes operational recovery with application logic.
The solution is the Supervisor pattern. By separating the worker (which performs the task) from the supervisor (which monitors the worker), you ensure that processes are automatically restarted to a known good state without crashing the entire application.
Prerequisites
- Erlang/OTP installed (Version 24+ recommended).
- Basic familiarity with the Erlang shell and module compilation.
- A worker module that exports a start function returning
{ok, Pid}.
Defining the Worker Specification
A supervisor requires a Child Specification—a declarative map that tells the supervisor how to start the process and when to give up on it. This prevents "infinite restart loops" where a process crashes immediately upon startup, consuming CPU and preventing other system functions from running.
-module(my_worker).
-export([start/0, loop/0]).
start() ->
spawn_link(fun my_worker/loop())
loop() ->
io:format("Worker is running...~n"),
receive
stop -> halt,
_ ->> loop()
end.
Configuring the Supervisor Strategy
The supervisor's behavior is defined by its Restart Strategy. Choosing the wrong strategy can cause unnecessary downtime for healthy processes.
| Strategy | Behavior | Best Use Case |
|---|---|---|
one_for_one |
Only the crashed process is restarted. | Independent workers (e.g., individual TCP connections). |
one_for_all |
All other children are terminated and then all are restarted. | Tightly coupled processes that depend on each other's state. |
rest_for_one |
Processes started after the crashed process are restarted. | Linear dependencies (Process B depends on Process A). |
Step-by-Step Implementation
Run these steps in your development environment to build a basic supervised hierarchy.
- Create the Supervisor Module: Define the
init/1function to specify the child list and the restart intensity (e.g., max 3 restarts every 5 seconds).
-module(my_supervisor).
-behaviour(supervisor).
-export([start_link/0, init/1]).
start_link() ->
supervisor:start_link({local, name, my_sup}, ?MODULE, init, []).
init([]) ->
ChildSpecs = [
#{id => worker_1,
start => {my_worker, start, []},
restart => permanent,
shutdown => 5000,
type => worker}
],
{ok, {#{strategy => one_for_one, intensity => 3, period => 5}, ChildSpecs}}.
- Compile and Start: Run
c(my_worker).andc(my_supervisor).in the Erlang shell, then executemy_supervisor:start_link(). - Verify Recovery: Find the PID of the worker and force it to crash.
Command to run (Erlang Shell):
Pid = whereis(worker_1). (Assuming the worker registered itself)
exit(Pid, kill).
Expected Check: The supervisor will detect the exit signal. Because the strategy is one_for_one and the restart is permanent, a new PID will be spawned immediately. You can verify this by checking the process list or using observer:start() to see the new node in the tree.
Critical Limitations and Risks
- State Loss: Supervisors restart processes from their initial state. Any data held in the worker's memory at the time of the crash is lost. To persist state, use an external database or an
etstable. - Circular Dependencies: Never allow a supervisor to be a child of a process it is supervising. This creates a circular dependency that can crash the entire VM.
- Restart Intensity: If the
intensityis set too high, a failing process may flood the system with restart attempts. If set too low, a few transient network glitches may cause the supervisor itself to terminate, killing all healthy children.
Rollback and Recovery
If a supervisor enters a failure loop and shuts down the tree, the state change is the termination of all child processes. To recover:
- Identify the root cause of the worker crash via logs.
- Fix the code or external dependency.
- Manually restart the supervisor using
my_supervisor:start_link().
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.