Configuring One-for-One Supervision in Erlang OTP
Wrap long-lived processes in OTP supervisors to gain declarative restart policies, bounded intensity, and ordered shutdown without manual exit trapping.
09 Aug 2026, 11:20 UTC

In production Erlang systems, long-running processes inevitably encounter unexpected conditions, from message mismatches to external resource failures. Attempting to handle every crash through manual exit trapping often produces processes that are alive but inconsistent, or error-handling logic that obscures the root cause. The practical takeaway: declare your fault-tolerance policy with an OTP supervisor and let the framework manage restarts, intensity, and shutdown ordering.
One-for-One Supervision Mechanics
A one_for_one strategy isolates failures to the crashed child. When that child exits, the supervisor restarts only that instance, leaving sibling processes untouched. The policy is configured in init/1 via SupFlags and a child specification list.
Minimal supervisor using map-based child spec (OTP 24+)
-module(my_app_worker_sup).
-behaviour(supervisor).
export([start_link/0, init/1]).
start_link() ->
supervisor:start_link({local, ?MODULE}, init, []).
init([]) ->
SupFlags = #{
strategy => one_for_one,
intensity => 1,
period => 5
},
ChildSpecs = [
#{
id => worker_1,
start => {my_worker, start_link, []},
restart => permanent,
shutdown => 5000,
type => worker
}
],
{ok, {SupFlags, ChildSpecs}}
end.
- intensity: maximum restarts allowed within the period. If exceeded, the supervisor terminates, escalating the failure.
- period: time window in seconds for intensity calculation.
- restart => permanent: child is always restarted, regardless of normal or crash termination.
- shutdown: milliseconds the supervisor waits for graceful exit before killing the process.
Strategy Selection for Coupled Children
| Strategy | Behavior | Use Case |
|---|---|---|
| one_for_one | Only the crashed child restarts. | Independent workers (e.g., separate TCP connections). |
| one_for_all | All other children terminated and then all restarted. | Tight coupling where children share state or dependencies. |
| rest_for_one | Children started after the crashed child are terminated and restarted. | Linear dependencies (e.g., Child B depends on Child A). |
The restart loop (CPU exhaustion)
A common production failure occurs when a process crashes immediately upon startup (e.g., due to a missing config file or database outage). If the intensity is too high or the period too long, the system may enter a rapid crash-restart loop, consuming 100% CPU and flooding logs.
Solution: keep intensity low. It is better for the supervisor to crash and escalate the failure to a higher-level supervisor (which might have a longer period) than to loop indefinitely.
Confusing availability with durability
Supervision ensures availability (the process is running), not durability (data is preserved). When a supervisor restarts a child, the process state is wiped. If you need to persist state across crashes, use an external store or a separate process not restarted by the same failure event, such as an ETS table owned by a separate process or a database.
Over-using catch-all error handling
Avoid wrapping gen_server callbacks in massive try blocks. The supervisor is designed to handle badmatch or case_clause errors. Catching everything hides the failure from the supervisor, preventing the restart mechanism from triggering and potentially leaving the process in a corrupted state.
Verification steps
- Start the supervisor:
my_app_worker_sup:start_link(). - Identify the worker PID:
supervisor:which_children(my_app_worker_sup). - Force a crash by sending a malformed message or using
exit(Pid, kill). - Run
supervisor:which_children(my_app_worker_sup)again. A new PID for worker_1 should appear.
Risk note: Using exit(Pid, kill) bypasses graceful shutdown. Verify that your shutdown timeout in the child spec is sufficient for your process to close sockets or file handles.
Restarting a process resets its state; supervision gives availability, not state durability, so persist critical state in ETS owned by a separate process, a database, or reconstruct it on init.
A child that crashes immediately on start will hit the restart intensity cap and take the whole supervisor down; validate configuration and external dependencies early in init/1.
Version-sensitive: child spec maps and dynamic-child APIs are standard in modern OTP (roughly OTP 18+ for maps, OTP 24+ for current docs); very old codebases may still use the tuple-based supervisor:child_spec format.
Supervisors do not protect against infinite loops or memory growth, only crashes; add monitoring (e.g., recon-style checks or alarms) for hung processes.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.