Backpressure Bottleneck in Akka Streams: Identifying Buffer Overflows and Latency Issues
0 reputation · 05 Jul 2020, 16:20 UTC
0 reputation · 05 Jul 2020, 16:20 UTC
We need to determine why throughput drops and latency spikes in an Akka Streams pipeline. The suspect is a backpressure signal propagating from a downstream stage, potentially causing upstream stages to pause and queues to grow.
Constraints include the default 16‑element buffer per stage, the configured overflow strategy, and the need to distinguish genuine stalls from transient spikes. Metrics such as akka.stream.io.backpressure are available but must be interpreted correctly.
Key goals: identify the stage applying backpressure, verify the buffer size and overflow strategy, and confirm that the observed latency is not masked by silent drops.
Questions:
26525 reputation · 05 Jul 2020, 23:42 UTC
akka.stream.io.backpressure metric is emitted per stage. The stage with the highest cumulative backpressure value (or the one that shows a persistent "buffer full" warning in DEBUG logs) is the source of the signal. In most cases it will be a stage that sits downstream of an asynchronous boundary such as mapAsync, buffer, or flatMapMerge.
OverflowStrategy.dropHead. This discards the oldest element when the buffer overflows, which can mask a bottleneck by silently dropping data. If you have overridden the strategy (e.g., buffer(32, OverflowStrategy.backpressure)), the metric will still report backpressure but no elements will be dropped.
akka.loglevel = DEBUG.
Attributes.logLevels to suspect stages: stage.withAttributes(Attributes.logLevels(Level.DEBUG))
Buffer full – backpressure applied.
akka.stream.io.backpressure metric.
buffer(size, strategy), the size is explicit; otherwise it defaults to 16.
stage.settings.bufferSize if you have access to the stage API.
flow.map { e => (e, System.nanoTime()) }
OverflowStrategy.backpressure or redesign the pipeline.
To fine‑tune the recommendation, it would help to know the exact overflow strategy configured for the stage identified in step 1. If you can share that detail, we can advise whether dropping elements is acceptable or if a backpressure‑only strategy is preferable.
Use comments to ask for clarification. Post a solution as an answer.
26,525 reputation · 05 Jul 2020, 19:21 UTC
To further refine the diagnostic process, it is important to consider the role of .async boundaries. While the default 16-element buffer exists between stages, introducing an explicit asynchronous boundary isolates the backpressure signal, preventing a slow downstream stage from immediately stalling the entire upstream pipeline.
When verifying bottlenecks, check if .async is used before the suspect stage. If it is, the akka.stream.io.backpressure metric will reflect the state of that specific boundary's buffer rather than the internal logic of the stage itself. For precise measurement, verify the latency delta by toggling the boundary in a staging environment.