Structured concurrency scope bounds during zero-downtime service migration
0 reputation · 17 May 2022, 22:11 UTC
Context
A small Kotlin service is being migrated to a new version using a rolling deployment where old and new instances coexist briefly. The migration tasks run as suspend functions on Dispatchers.IO within a CoroutineScope tied to the service lifecycle component.
Goal
Ensure that all in-flight migration coroutines — including blocking I/O offloaded to Dispatchers.IO — terminate cleanly when the old instance receives a shutdown signal, without leaking resources or leaving orphaned operations on the shared thread pool.
Constraints and uncertainty
Structured concurrency guarantees that child coroutines complete or are cancelled before the parent scope finishes. However, the documentation does not explicitly clarify whether this guarantee extends across process boundaries during a rolling update, nor how CancellationException propagates into blocking calls already executing on Dispatchers.IO. Using GlobalScope is ruled out due to leak risk, but the lifecycle scope may be destroyed before the dispatcher finishes interruptible I/O.
Questions
- Does a parent
CoroutineScopecancellation reliably interrupt blocking operations running onDispatchers.IOwhen the process is terminating? - What is the expected behavior of child coroutines if the JVM shuts down before
Job.join()completes on the dispatcher threads?