Akka Cluster upgrade recovery: mixed-version tolerance or blue-green cutover?
0 reputation · 22 Feb 2022, 13:27 UTC
Akka Cluster documents rolling upgrades in which nodes of adjacent versions coexist while akka.cluster.allow-weakly-upgraded-nodes is enabled. That tolerance assumes message and event serialization remains backward compatible, typically through a schema-evolution format rather than Java serialization. The alternative is orchestration-driven replacement, such as a Kubernetes blue-green cutover, where no mixed-version window exists but the zero-downtime constraint shifts onto traffic draining and membership shutdown.
The unresolved decision concerns recovery rather than the upgrade itself. If a node fails mid-upgrade and cannot rejoin, the configured downing provider decides when it is marked down. Tolerating a mixed-version remainder and abandoning it for a fresh cluster imply different failure modes, and the documentation does not settle which is safer under a strict zero-downtime requirement.
- Which documented approach bounds recovery failure modes more predictably when a node cannot rejoin a partially upgraded cluster?
- What must hold for persistence event schemas and actor protocols before either approach is treated as safe?
- How should downing behavior be reasoned about while the cluster is intentionally mixed-version?
Because configuration keys and defaults differ across Akka releases, the trade-off should be confirmed against the documentation for the deployed version.