Diagnosing RabbitMQ Publisher Confirm Failures That Cause Publishing Stalls
Identify why RabbitMQ publisher confirms stall, diagnose queue, memory, network, or client issues, and apply targeted fixes with verification steps.
04 Dec 2025, 23:15 UTC

Recognizable condition
Publishers report increased latency or timeouts when sending messages. The broker log contains messages such as "channel error: precondition failed" or "no publishers confirm" for the affected connections.
Cause / diagnostic table
| Possible cause | What to look for |
|---|---|
| Full queue or high memory usage | Queue length growing steadily; node memory approaching the high‑watermark. |
| Network partition preventing confirms from reaching the client | Dropped or retransmitted TCP packets; connection state shows "CLOSE_WAIT" or frequent reconnects. |
| Client library not handling confirm callbacks | Publish calls return immediately but no confirm/ack is processed; exceptions are swallowed. |
Ordered checks
- Verify queue length and memory usage
Run on any RabbitMQ node (requires
rabbitmqctladmin rights):rabbitmqctl list_queues name messages memoryLook for queues with rapidly increasing
messagesormemoryvalues approaching the node’s RAM limit.Risk: None; read‑only.
- Check node health
On the same node:
rabbitmqctl statusExamine the
mem_limitandmem_usedfields; also note anyalarms(e.g., memory_alarm). - Inspect TCP connections for drops
From the client host or the broker:
netstat -anp | grep :5672or via RabbitMQ:
rabbitmqctl list_connections name peer_host stateFrequent transitions to
closingorclosedstates indicate network instability. - Confirm client library handling
Enable debug logging in the publisher (e.g., for Java client:
com.rabbitmq.clientlevel TRACE). Verify that:- The
ConfirmListener.handleAckorhandleNackcallback is invoked for each publish. - Exceptions from the callback are not caught and ignored.
If the library version is outdated, note it for upgrade.
- The
Fixes tied to findings
- High queue length / memory pressure
- Increase the queue length limit (if using a classic queue):
rabbitmqctl set_policy HA-queue '^my.queue' '{"max-length":100000}' --apply-to queuesOr add more consumers to drain the queue.
- If memory is the bottleneck, raise the VM memory high watermark (requires node restart):
# edit /etc/rabbitmq/rabbitmq.conf vm_memory_high_watermark = 0.7Then restart the node:
systemctl restart rabbitmq-server.Rollback: Restore the previous
vm_memory_high_watermarkvalue and restart. - Network partition
- Restore connectivity between client and broker (check firewalls, VPN, load balancer).
- Ensure mirrored queues are symmetric across nodes (
ha-mode: all).
After fixing, monitor
rabbitmqctl list_connectionsfor stablerunningstate.Rollback: If you temporarily disabled a firewall rule, re‑enable it after verification.
- Client mis‑handles confirms
- Upgrade to a recent client library version that properly exposes confirm callbacks.
- Implement retry/back‑off logic on
nackor timeout:
// pseudo‑code if (!confirmReceivedWithin(timeout)) { backoffAndRetry(); }Rollback: Revert to the previous library version if the new version introduces regressions.
Escalation criteria
- Publish latency remains > 2× baseline after applying the above fixes.
- Broker logs continue to show
precondition failederrors despite healthy queues and connections. - Memory alarms persist even after increasing
vm_memory_high_watermarkand adding consumers.
In these cases, consider:
- Opening a support ticket with RabbitMQ vendor, providing
rabbitmqctl reportoutput. - Reviewing cluster topology for split‑brain scenarios.
- Conducting a controlled load test in a staging environment to isolate the issue.
Verification
After applying fixes:
- Monitor publish latency via
rabbitmq_exporter(Prometheus metricrabbitmq_queue_messages_publish_rate_totalor latency histograms). Latency should return to pre‑incident baseline (e.g., <200 ms for 95th percentile). - Check broker logs for absence of
channel error: precondition failedandno publishers confirmmessages over a 10‑minute window. - Run a synthetic publisher that sends 1 000 messages with confirms enabled and asserts each publish receives a confirm within the configured timeout (
<200 ms) for a sustained period (e.g., 5 min). No missed confirms should be reported.
Limitations
- Queue length checks may not detect transient spikes that resolve before the sampling interval.
- Network partition detection relies on TCP state; application‑level partitions (e.g., client‑side throttling) may not appear.
- Changing
vm_memory_high_watermarkrequires a node restart, which briefly impacts availability.
Practical way to check the result: after any change, re‑run the ordered checks and confirm that the abnormal indicators (queue growth, memory usage, connection drops, missing confirm callbacks) are no longer present.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.