Architecting Fault-Tolerant Services with Gleam OTP Supervisors
Learn how to use Gleam OTP Supervisors to build fault-tolerant services using the 'let it crash' philosophy, process isolation, and restart strategies.
03 Feb 2026, 11:09 UTC

The Problem: Preventing Total Service Collapse
In concurrent systems, a single unhandled exception in a worker process can lead to a cascading failure if that process is critical to the system's availability. The goal is to isolate failures so that a crashing worker does not take down the entire application, and is instead restored to a known good state automatically.
The takeaway is to use a Supervisor—a specialized process that monitors other processes (children)—to implement a "let it crash" philosophy. By decoupling the failure detection from the business logic, you ensure the service remains available even during transient errors.
The Smallest Suitable Design
For a basic fault-tolerant service, the leanest architecture consists of a Supervisor managing a single GenServer (Generic Server). The GenServer handles the state and request logic, while the Supervisor manages the lifecycle.
One-for-One Strategy
The most common strategy for independent services is one_for_one. In this configuration, if a child process terminates, only that process is restarted. Other children are left untouched. This is ideal for services where workers do not depend on each other's internal state.
Implementation Example
To implement this, you must add gleam_otp to your gleam.toml. The following conceptual configuration defines a supervisor that manages a worker process:
// Note: This example assumes gleam/otp is installed
import gleam/otp/supervisor
import gleam/otp/genserver
// The child specification tells the supervisor how to start the worker
let child_spec = supervisor.child_spec(
start_spec: genserver.start_spec(my_worker_init),
restart_strategy: supervisor.OneForOne
)
// Start the supervisor with a list of children
supervisor.start(supervisor.Config({
children: [child_spec],
max_restarts: 3,
max_seconds: 5
}))
Trust and Data Boundaries
Gleam leverages the BEAM (Erlang Virtual Machine) process model, which creates strict boundaries between concurrent entities:
- Heap Isolation: Each process owns its own memory heap. A crash in the GenServer cannot corrupt the memory of the Supervisor.
- Typed Mailboxes: Communication happens via asynchronous message passing. The trust boundary is the mailbox interface; the Supervisor does not share mutable state with the worker, preventing race conditions during restarts.
- Immutable Messages: Because data passed between processes is immutable, the Supervisor can safely restart a worker without worrying about partially modified shared objects.
Operational Checks and Failure Modes
A Supervisor is not a magic bullet; it requires tuning to avoid "restart loops" where a process crashes immediately upon booting, consuming CPU cycles indefinitely.
Restart Thresholds
Supervisors use max_restarts and max_seconds to detect permanent failures. For example, if configured for 3 restarts in 5 seconds, the Supervisor will attempt to revive the worker three times. If the worker crashes a fourth time within that window, the Supervisor concludes the failure is not transient.
Failure Escalation
When the restart threshold is exceeded, the Supervisor itself terminates. This is a deliberate design choice called escalation. The failure is pushed up to the next level of the hierarchy (e.g., a top-level Application Supervisor), which may then decide to restart the entire subsystem or shut down the node to prevent inconsistent state.
| Failure Type | Behavior | Outcome |
|---|---|---|
| Transient (e.g., Timeout) | Worker crashes $\rightarrow$ Supervisor restarts | Service restored in milliseconds |
| Permanent (e.g., Bad Config) | Worker crashes $\rightarrow$ Threshold reached | Supervisor terminates (Escalation) |
| Systemic (e.g., OOM) | VM crashes | External orchestrator (K8s/Systemd) restarts node |
Constraints and Design Shifts
This architecture is suitable for localized state management, but certain conditions require a change in design:
- Blocking I/O: If a GenServer performs long-running synchronous I/O, it blocks the process scheduler. This can delay the Supervisor's ability to monitor the process, leading to false positives or unresponsive services. In these cases, move I/O to
gleam/taskor Erlang ports. - Distributed State: A single Supervisor is limited to one node. If you need to share state across a cluster, you must integrate a global registry or use
:pg(Process Groups) via Gleam's FFI to the Erlang runtime. - Complex Dependencies: If Worker B cannot function without Worker A, the
one_for_onestrategy must be replaced withone_for_all, ensuring that if A crashes, B is also restarted to reset the dependency chain.
Verification Steps
To verify the fault tolerance of your implementation:
- Run the application using
gleam run -m your_app/main. - Trigger a controlled crash in the worker (e.g., using a specific "kill" message that calls
exit()). - Observe the logs to confirm the Supervisor detected the termination and spawned a new PID.
- Verify that the service continues to respond to requests after the restart.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.