Solving Message Loss in NATS with JetStream Persistence
Stop losing messages when subscribers go offline. Learn how to use NATS JetStream to implement durable persistence and at-least-once delivery guarantees.
08 May 2026, 00:53 UTC

The 'Fire-and-Forget' Problem
Core NATS is an incredibly fast messaging system, but by default, it operates on a fire-and-forget model. If a subscriber is offline when a message is published, that message is gone forever. For telemetry or real-time dashboards, this is acceptable. For order processing, billing, or critical state updates, it is a deal-breaker.
The solution is JetStream. JetStream evolves NATS from a simple message bus into a distributed log, allowing you to persist messages on disk and guarantee that a consumer eventually receives them, regardless of network partitions or application crashes.
How JetStream Ensures Delivery
JetStream introduces the concept of a Stream, which is a configured storage layer that captures messages published to specific subjects. Unlike core NATS, where the server doesn't track who received what, JetStream tracks the state of every message relative to a Consumer.
A Consumer is essentially a view into a stream. It tracks which messages have been delivered and acknowledged. By using At-Least-Once delivery, NATS will continue to redeliver a message until the client explicitly sends an acknowledgement (Ack). If the client crashes mid-process, the message remains in the stream and will be redelivered to the next available instance of that consumer.
Choosing the Right Consumer Strategy
When implementing JetStream, the most critical decision is choosing between Push and Pull consumers. While Push consumers send messages to the client as they arrive, Pull Consumers allow the client to request a specific batch of messages when it has the capacity to process them.
Pull consumers are generally preferred for production workloads because they prevent the "poison pill" scenario where a slow consumer is overwhelmed by a flood of messages, leading to memory exhaustion or cascading failures.
Comparison: Ack Policies
| Policy | Behavior | Best Use Case |
|---|---|---|
AckExplicit |
Every single message must be acknowledged individually. | Critical financial transactions. |
AckAll |
Acknowledging message N acknowledges all messages before it. | Sequential logs or telemetry streams. |
AckNone |
Messages are considered delivered the moment they are sent. | High-throughput, non-critical data. |
Worked Example: Creating a Durable Stream
To enable JetStream, you must start the NATS server with the -js flag. The following example uses the nats CLI to set up a durable stream for an order processing system.
1. Create the Stream
Run this command from your terminal with the NATS CLI installed. This creates a stream named ORDERS that listens to any subject starting with orders.
nats stream add ORDERS --subjects "orders.*" --storage file --retention limits --max-msgs=-1 --max-bytes=-1
Risk: Setting --max-msgs=-1 and --max-bytes=-1 disables limits. In a production environment, always set a byte limit to prevent disk exhaustion on the NATS server.
2. Create a Pull Consumer
Create a durable consumer named order-processor. A durable consumer remembers its progress even if the client disconnects.
nats consumer add ORDERS order-processor --pull --ack explicit --deliver all
3. Verification
Publish a test message and attempt to fetch it using the CLI:
nats pub orders.new "Order #1234"
nats consumer next ORDERS order-processor
If you run the next command without acknowledging the message (or if the process fails), running it again will return the same message, verifying the at-least-once guarantee.
Trade-offs and Limitations
Persistence comes with a performance cost. Writing to disk is orders of magnitude slower than memory-only routing. Additionally, JetStream requires a more complex clustering setup (using Raft) to ensure that data is replicated across multiple nodes. If you only run a single NATS node, you have persistence but no high availability; if that node's disk fails, your data is lost.
Another limitation is message duplication. Because JetStream guarantees at-least-once delivery, it is possible for a consumer to process a message, crash before sending the Ack, and then receive the same message again upon restart. Your application logic must be idempotent (processing the same message twice should have no additional effect).
Closing Action
To migrate from core NATS to JetStream, start by identifying your most critical data paths. Enable the -js flag on your server, define a stream with a strict max-bytes limit to protect your infrastructure, and implement a Pull Consumer with AckExplicit to ensure no message is dropped during processing.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.