Diagnosing and Fixing Consumer Lag in NATS JetStream
A step‑by‑step diagnostic guide for detecting and resolving consumer lag in NATS JetStream, covering metrics, stream limits, consumer settings, and server resources. Includes concrete examples, escalation criteria, and monitoring best practices.
05 Jan 2026, 12:03 UTC

Problem Statement
Consumer lag in JetStream appears when the Delivered count of a consumer grows faster than the Acked count. The consumer is behind the stream head and may be replaying many messages, which can overwhelm application logic or lead to missed deadlines.
Understanding Lag Metrics
Run nats consumer info <stream> <consumer> on the node where the consumer runs. Key fields:
- Delivered – total messages sent to the consumer.
- Acked – total messages the consumer has acknowledged.
- Pending – messages queued for delivery (Delivered – Acked).
- Last Delivered – sequence number of the last delivered message.
- Last Acked – sequence number of the last acknowledged message.
Lag exists when Delivered > Acked and Pending is non‑zero.
Diagnostic Checklist
- Check consumer configuration
nats consumer info <stream> <consumer> | grep -E 'Ack Wait|Max Deliver|Replay' - Inspect stream limits
nats stream info <stream> | grep -E 'max_age|max_bytes|max_msgs' - Verify server resource usage
curl http://localhost:8222/metrics | grep -E 'jetstream_disk|go_memstats|process_cpu_seconds_total' - Confirm network health
nats ping - Detect paused consumers
nats consumer info <stream> <consumer> | grep Paused
Common Causes & Fixes
| Cause | Diagnostic Indicator | Fix |
|---|---|---|
| Consumer processing slower than publish rate | High Pending count, Delivered growing rapidly | Increase consumer CPU, add more consumer instances, or throttle publisher. |
| Older messages purged before ack | Delivered jumps ahead of Last Acked, max_age or max_bytes low | Raise retention values or switch to max_msg based retention. |
| Repeated deliveries, ack timeouts | Many redeliveries in consumer logs, ack_wait < 2s | Increase ack_wait to 5–10s or adjust max_deliver. |
| Disk I/O, memory, or CPU saturation | High metrics from Prometheus, slow JetStream persistence | Upgrade storage, increase RAM, or add nodes to cluster. |
| Consumer paused or marked slow | Consumer status shows Paused or Slow | Resume consumer or investigate application backpressure. |
| Consumer disconnects, backlog grows on reconnection | Cluster health shows split brain, consumer reconnect logs | Repair network, ensure quorum, or use cluster replication. |
| Immediate replay overwhelms consumer | Sudden spike in Pending after consumer start | Use Replay=All with a back‑off, or split backlog into smaller batches. |
| No new consumers, existing ones slow | Stream info shows max_consumers reached | Increase max_consumers or delete stale consumers. |
Concrete Example
Suppose a consumer reports:
Delivered: 1,200,000
Acked: 950,000
Pending: 250,000
Run the diagnostics:
- Check stream limits:
Thenats stream info orders | grep -E 'max_age|max_bytes|max_msgs' # Output: # Max Age: 1h # Max Bytes: 500MB # Max Messages: 1,000,000max_msgsis lower than Delivered, so older messages are purged before the consumer can ack. - Adjust retention:
After the update, monitor Pending until it drops.nats stream update orders --max-msgs 2_000_000 - If Pending remains high, check server metrics. If disk I/O is 90%, consider adding a fast SSD or scaling the cluster.
Escalation Criteria
- If
Pendingexceeds 10% ofDeliveredfor > 5 minutes. - When server metrics show sustained CPU > 80% or memory > 75%.
- Network partitions detected for > 30 seconds.
- Consumer fails to resume after 3 reconnection attempts.
Escalate to the platform ops team, provide the diagnostic logs, and request a capacity review or cluster upgrade.
Monitoring & Long‑Term Prevention
- Set Prometheus alerts on
jetstream_consumer_pendingandjetstream_stream_retentionlimits. - Use the
nats monitorcommand to watch consumer health in real time. - Periodically review
ack_waitandmax_deliversettings against observed message rates. - Implement back‑pressure in the consumer application to pause publishing when the backlog grows beyond a threshold.
Limitations
JetStream behavior changes between major releases; always validate configuration parameters against your deployed NATS version. Network latency can temporarily inflate perceived lag; confirm connectivity before attributing lag to consumer logic.
Practical Check
After applying a fix, run again:
nats consumer info orders my_consumer | grep Pending
# Expect Pending to drop below 5% of Delivered
If the condition persists, revisit the diagnostic checklist.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.