Beyond Try-Catch: Implementing Fault Tolerance with Erlang Supervision Trees
Stop trying to catch every exception. Learn how Erlang's 'Let it Crash' philosophy and Supervision Trees provide a robust framework for automatic recovery and high availability.
19 Jul 2026, 00:28 UTC

The Cost of Defensive Programming
In most languages, the instinct is to wrap every risky operation in a try-catch block. This defensive approach assumes that if you can anticipate every possible failure—network timeouts, null pointers, or corrupted data—you can keep the system running. However, this often leads to "zombie states," where a process survives a crash but remains in an inconsistent or corrupted internal state, leading to unpredictable bugs later.
Erlang takes the opposite approach: Let it Crash. Instead of attempting to handle every edge case within the business logic, Erlang isolates tasks into lightweight processes that share no memory. When a process hits an unrecoverable state, it terminates immediately. The goal is not to prevent the crash, but to ensure that a separate, healthy process (the Supervisor) can restore the system to a known-good state.
The Anatomy of a Supervision Tree
A supervision tree is a hierarchy of processes where "worker" processes perform the actual tasks and "supervisor" processes manage their lifecycles. This separation of concerns ensures that the logic for doing the work is decoupled from the logic for recovering from failure.
Supervisors monitor their children via links. If a child process terminates abnormally, the supervisor receives a signal and applies a restart strategy to resolve the issue. This creates a failure domain; a crash in a low-level worker doesn't necessarily bring down the entire application, only the specific branch of the tree responsible for that task.
Choosing a Restart Strategy
Not all failures are independent. Depending on how your processes rely on one another, you must choose a specific OTP (Open Telecom Platform) restart strategy:
- one_for_one: If a process crashes, only that process is restarted. This is ideal for independent workers, such as individual client connections.
- one_for_all: If one process crashes, the supervisor terminates and restarts all other processes it is managing. Use this when a group of processes are tightly coupled and cannot function if one member is missing.
- rest_for_one: If a process crashes, only that process and any processes started after it in the sequence are restarted. This is useful for linear dependencies (e.g., a database connection must be active before a cache manager starts).
Worked Example: A Simple Supervised Worker
To implement this, you typically use the supervisor behavior. In this example, we define a worker that crashes when it receives a specific "poison pill" message, and a supervisor that ensures it stays alive.
% Run this in the Erlang shell (erl) with OTP installed
% 1. Define a simple worker module
-module(my_worker).
-export([start/0, loop/0]).
start() ->
spawn(my_worker, loop, []).
loop() ->
receive
{crash, _} ->
exit(forced_crash) % Simulate a fatal error
{msg, Content} ->
io:format("Worker received: ~p~n", [Content]),
loop()
end.
% 2. Define the Supervisor configuration
% In a real app, this is handled by the supervisor behavior module
% Strategy: one_for_one, Max Restarts: 3, Time Window: 5 seconds
% Child Spec: {id, start_function, {args}, restart_type}
Verification: To test this in a live environment, you would start the supervisor and send a {msg, "Hello"} to the worker. Then, send {crash, true}. You will observe the worker process ID (PID) change in the logs, indicating the supervisor has detected the death and spawned a fresh instance of the worker.
The Danger of the Infinite Restart Loop
The "Let it Crash" philosophy has a critical failure mode: the permanent bug. If a process crashes because of a malformed configuration file or a specific piece of bad input, restarting it will simply lead to another crash. This creates a rapid loop of crash → restart → crash.
To prevent this from consuming all system resources, supervisors use a restart intensity limit (e.g., MaxRestarts). If a supervisor exceeds the allowed number of restarts within a specific time window, the supervisor itself will crash. This propagates the failure up the tree to a higher-level supervisor, which may then decide to restart the entire subsystem or shut down the node to prevent a cascading failure.
Practical Implementation Checklist
- Isolate State: Ensure workers do not share memory. Use messages to communicate.
- Define Dependencies: Use
one_for_allonly if processes are functionally interdependent. - Tune Intensity: Set
MaxRestartsbased on your expected transient failure rate to avoid premature system shutdowns. - Log the Crash: Use a crash reporter to capture the reason for the exit before the process is wiped from memory.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.