Akka OneForOne Supervision for Fault‑Tolerant Bounded‑Context Microservices
Describes how a single top‑level OneForOneSupervisor can isolate actor failures in a bounded‑context microservice while preserving message ordering, with configuration, operational checks, and conditions that would require a redesign.
15 Jan 2026, 16:09 UTC

Requirements
The microservice must keep running even if an individual actor throws an exception. A crash in one actor should not bring down the whole ActorSystem, but messages addressed to that actor must still be processed in order after recovery. The supervision mechanism therefore needs to isolate failures while preserving per‑actor message ordering.
Smallest suitable design
A single top‑level supervisor per microservice using OneForOneStrategy is sufficient. The supervisor watches only the top‑level business actors that own distinct state partitions (e.g., one actor per aggregate root). All other actors are created as children of those business actors and inherit the same supervision directive.
// application.conf
akka {
actor {
deployment {
/businessActor-* {
supervisor-strategy = "akka.supervision.OneForOneStrategy"
}
}
}
}
In code the supervisor is defined as:
import akka.actor.{OneForOneStrategy, Props, SupervisorStrategy, Actor, ActorRef}
import scala.concurrent.duration.
class ServiceSupervisor extends Actor {
override val supervisorStrategy = OneForOneStrategy(
maxNrOfRetries = 10,
withinTimeRange = 1.minute,
loggingEnabled = true
) {
case _: Exception => SupervisorStrategy.Restart
}
// create business actors, each owning a partition of state
val orderActor: ActorRef = context.actorOf(Props[OrderActor], "orderActor")
val inventoryActor: ActorRef = context.actorOf(Props[InventoryActor], "inventoryActor")
def receive: Actor.Receive = {
case msg => // forward to appropriate business actor
}
}
Trust and data boundary
Messages that cross actor boundaries are considered untrusted. The receiving actor must validate the payload type and content before using it. Unexpected or malformed messages are treated as poison pills: the actor stops processing them and lets the supervisor apply its directive (usually a restart). This prevents corrupt state from propagating.
def receive: Actor.Receive = {
case cmd: ValidCommand => process(cmd)
case unknown =>
log.warning("Received unexpected message type {}; treating as poison pill", unknown.getClass)
context.stop(self) // supervisor will restart
}
Operational checks
- JMX monitoring: Enable
akka.jmx.enabled=onand observe theSupervisorRestartCountmetric for the top‑level supervisor. Each restart increments the counter. - DeathWatch: The supervisor calls
context.watch(child)on each business actor. When a child terminates, aTerminatedmessage is received; logging this event helps detect supervision chain breaches. - Mailbox size: Configure a bounded mailbox (e.g.,
mailbox-capacity = 1000) and monitorakka.actor.mailbox.queue-sizevia JMX or Lightbend Telemetry. A steadily growing queue signals that the supervisor is not restarting fast enough. - Log escalation: If the supervisor escalates a failure (because the directive does not match), the event appears at
ERRORlevel with the pattern "Escalating failure". Alert on these logs to catch mis‑configured strategies.
Failure modes and design‑change conditions
If the supervisor’s maxNrOfRetries is exceeded within the withinTimeRange, it escalates the failure to its own supervisor (the ActorSystem guardian). Repeated escalation causes the ActorSystem to shut down. This is a clear signal that the current supervision design is insufficient.
Conditions that would trigger a redesign:
- Persistent state: When business actors rely on durable storage (e.g., event‑sourced actors), a simple restart may lose in‑memory buffers. Switch to a
PersistentSupervisorthat persists supervision directives or uses Akka Persistence’sReceiveRecoverto rebuild state after restart. - Remote deployment: If actors are deployed on remote nodes, a local supervisor cannot restart a crashed remote child. Use a
RemoteSupervisorpattern where each node runs its own supervisor and monitors remote children via DeathWatch. - Shared mutable state: When siblings share state that must be recovered together,
OneForOneStrategyis too narrow. Replace it withAllForOneStrategy(or a custom strategy) to restart the entire sibling set, accepting the extra churn.
Limitations and practical verification
The OneForOneStrategy only restarts the failed actor; it does not protect against cascading failures caused by exhausted resources (e.g., thread starvation) or by messages that repeatedly cause the same exception. To verify the design in a safe environment:
- Create a test ActorSystem with Akka TestKit.
- Inject a throwing exception in a child actor’s
receivemethod. - Assert that the supervisor logs a restart event and that the child resumes processing subsequent messages without the ActorSystem terminating.
- Enable JMX (
akka.jmx.enabled=on) and confirm that theSupervisorRestartCountmetric increments after each failure. - Run a modest load generator (e.g., Gatling) and watch the mailbox size metric; it should remain within the configured bound and not show uncontrolled growth.
If any of these checks fail—particularly if the ActorSystem shuts down or the restart count stops increasing—revisit the supervision directive, retry limits, or consider moving to a persistent or remote supervision model.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.