Moving Beyond Try-Catch: Implementing Erlang's 'Let it Crash' Philosophy
Stop fighting runtime errors with nested try-catch blocks. Learn how Erlang’s fault tolerance patterns let you build self‑healing systems using supervisors and the "Let it Crash" philosophy.
20 Nov 2025, 23:25 UTC

The Fragility of Defensive Programming
In most languages, a runtime exception is a crisis. Developers spend significant effort writing deeply nested try-catch blocks to anticipate every possible edge case – null pointers, timeouts, malformed packets. This defensive style often leads to a "zombie" state where a process survives a crash but remains inconsistent, causing bugs later.
Erlang solves this by treating failure as a first‑class citizen. Instead of preventing every error, Erlang encourages a Let it Crash philosophy: allow a process to terminate when it hits an unexpected state, and let a supervisor bring the system back to a known‑good configuration.
Isolation via Lightweight Processes
The foundation of this approach is process isolation. Erlang processes are not OS threads; they are lightweight entities with their own private heaps and stacks. Because processes share no memory, a crash in one cannot corrupt another’s memory.
When a process crashes, it sends an exit signal to any linked processes. The supervisor – a specialized process – monitors its children and restarts them according to a strategy.
Choosing a Supervision Strategy
- one_for_one: If a child crashes, only that child is restarted. Ideal for independent workers.
- one_for_all: If one child crashes, the supervisor terminates and restarts all children. Use when processes are tightly coupled.
- rest_for_one: If a child crashes, that child and any started after it are restarted. Useful for linear dependencies.
Worked Example: A Basic Supervisor
Define a worker module and a supervisor that manages it. When the worker is killed, the supervisor automatically respawns it.
% Run this in the Erlang shell (erl)
% Worker module
-module(worker).
-export([start/0, loop/0]).
start() -> spawn(worker, loop, []).
loop() ->
io:format("Worker is running...~n"),
timer:sleep(5000),
loop().
% Supervisor module
-module(my_sup).
-export([start_link/0]).
start_link() ->
ChildSpecs = [
#{id => worker_1,
start => {worker, start, []},
restart => permanent,
shutdown => 5000}
],
supervisor:start_link({one_for_one, 1, 5}, {local, my_supervisor}, ChildSpecs).
Verification and Testing
- Start the supervisor:
my_sup:start_link(). - Find the worker PID:
whereis(my_supervisor)and inspect children or usesupervisor:which_children(my_supervisor). - Force a crash:
exit(WorkerPid, kill). - Observe the logs: the "Worker is running..." message reappears almost immediately, indicating a restart.
Trade‑offs of Automatic Recovery
- Infinite restart loops: If a child fails immediately on startup, the supervisor will keep restarting it until the max restart intensity is exceeded. Then the supervisor itself crashes, escalating the failure.
- State loss: A crashed process loses all transient state. Persist critical data in an external store (Mnesia, PostgreSQL, Redis) if it must survive restarts.
- Hidden systemic bugs: Frequent restarts can mask problems. Ensure logging and monitoring capture the reasons for failure so they can be addressed.
Practical Implementation Path
Identify boundaries of failure in your application. Split logic into small, isolated processes, and assign each a supervisor. Use one_for_one for independent workers and one_for_all only when strict coupling exists. Add comprehensive logging to surface recurring crash reasons, and test restart behavior by manually killing child processes in a staging environment.
Conclusion
"Let it Crash" is not a shortcut to skip debugging; it is a design principle that leverages Erlang’s lightweight processes and supervision trees to build resilient systems. By isolating failures and allowing supervisors to recover automatically, you can achieve high uptime with minimal manual intervention.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.