Blue-green queue migration in RabbitMQ: drain in place or shovel to the new topology
0 reputation · 26 Jan 2024, 09:25 UTC
0 reputation · 26 Jan 2024, 09:25 UTC
We need to change immutable queue arguments (for example x-queue-type and max-length) on a set of classic queues in production. Since redeclaring with different arguments raises 406 PRECONDITION_FAILED, the plan is a blue-green approach: declare new queues and bindings alongside the old ones, move consumers over, then remove the old topology.
The unresolved decision is how to handle messages still sitting in the old queues at cutover time. One option is to keep consumers bound to both old and new queues until the old ones drain to zero. The other is to use the shovel plugin to move remaining messages into the new queues, accepting that ordering across the two topologies is not preserved and that duplicates are possible if acks are lost mid-transfer.
Messages are roughly order-sensitive per queue, and some consumers are not fully idempotent today. We are on a recent 3.x release and can verify behavior in a staging cluster first.
29275 reputation · 26 Jan 2024, 09:52 UTC
For non-idempotent consumers and order-sensitive messages, dual-binding consumers (drain-in-place) is generally safer than shoveling, provided you accept a temporary window of out-of-order processing across the two topologies.
The primary risk in your scenario is the lack of consumer idempotency. Here is how the two strategies impact your constraints:
| Factor | Dual-Binding (Drain) | Shoveling |
|---|---|---|
| Ordering | Preserved per-queue; broken across topologies. | Preserved per-message; broken across topologies. |
| Duplicates | Low (standard RabbitMQ ack risks). | Higher (risk of duplication during transfer/ack failure). |
| Complexity | Low (consumer config change). | Medium (plugin config and monitoring). |
Dual-binding does not introduce new ordering hazards within a single queue, but it does mean a message produced to the new queue might be processed before a message produced to the old queue. If your business logic requires strict global ordering across the migration boundary, neither method is sufficient without a sequencing key at the application level. However, it avoids the double-handling and potential duplication inherent in moving messages between queues.
The Shovel plugin preserves the order of messages as they are read from the source and written to the destination. In practice, for a single source-to-target pair, the sequence is maintained. However, because you are moving messages into a queue that is already receiving live traffic from producers, the "shoveled" messages will appear at the end of the new queue, effectively jumping behind any messages already produced to the new topology.
There is no single "drained" event signal. You must rely on the messages and messages_ready metrics. A queue is considered fully drained when:
messages count is 0.messages_unacknowledged is 0 (ensuring no consumer is still holding a message).rabbitmqctl list_queues name messages messages_readymessages and messages_unacknowledged reach 0, remove the Blue queue and its bindings.Missing Diagnostic: Are your consumers using a single-threaded prefetch (basic.qos prefetch_count=1)? If prefetch is high, the "drain" phase may take longer as messages remain unacknowledged in consumer buffers despite the queue appearing empty.
Use comments to ask for clarification. Post a solution as an answer.
No question comments on this page.