JetStream Pull Consumers: Let Your Workers Set the Pace
JetStream pull consumers let subscribers fetch batches at their own pace with explicit acks, giving you backpressure, at-least-once delivery, and queue-style scaling without extra infrastructure.
21 Jul 2025, 04:59 UTC

Push-based messaging has a quiet failure mode: the server decides how fast your consumers work. When a downstream database slows or a deploy cuts your worker count in half, messages pile into sockets, buffers fill, and you start dropping or timing out work you never asked to receive that fast. NATS JetStream's pull consumers invert that relationship — the subscriber asks for work when it's ready, processes it, and acknowledges it explicitly. That single inversion buys you backpressure, at-least-once delivery, and trivial horizontal scaling without extra infrastructure.
What a pull consumer actually is
JetStream is the persistence layer built into the NATS server. Two objects matter here: a stream stores messages published to a subject, and a consumer defines how a client reads from that stream. Consumers come in two flavors. Push consumers have the server deliver messages to a subject as fast as it can. Pull consumers sit idle until a client explicitly requests a batch — "give me up to 50 messages, wait up to 5 seconds if none are ready."
The client then processes the batch and sends an acknowledgement (an ack) for each message. If a message isn't acked within the consumer's AckWait window, JetStream redelivers it. That's the whole contract, and it's what makes the rest of the benefits fall out naturally.
Why pulling beats pushing for worker workloads
Backpressure for free. A slow consumer simply fetches less often or requests smaller batches. Messages accumulate safely in the stream — bounded by your retention and storage limits — instead of overwhelming a subscriber's inbox. Your processing rate and your arrival rate are decoupled.
Horizontal scaling without coordination. Multiple processes bound to the same durable pull consumer share the work queue-style: each message goes to exactly one member. Scaling out is literally starting more processes. There's no partition assignment, no rebalancing protocol, no leader election.
At-least-once semantics you can reason about. Crash mid-batch without acking, and those messages come back after AckWait expires. Combined with MaxDeliver (a cap on redelivery attempts), you get retries with a ceiling, and a place to route poison messages.
A worked example with the nats CLI
You can observe the entire lifecycle locally. Run these in a shell; the server needs no special permissions, and the CLI commands talk to it over the default localhost port.
# Terminal 1: start a server with JetStream enabled
nats-server -js
# Terminal 2: create a stream capturing the "jobs" subject
nats stream add JOBS --subjects "jobs" --retention work --storage file
# Create a durable pull consumer with a 30s ack window
nats consumer add JOBS WORKER --pull --deliver all --ack explicit --wait 30s --max-deliver 5
# Publish some work
nats pub jobs "resize-image-42"
nats pub jobs "resize-image-43"
# Fetch a batch of up to 10 messages
nats consumer next JOBS WORKER --count 10 --ackThe --retention work flag selects work-queue retention: a message is deleted from the stream once any consumer acks it, which suits task queues. For an event log you want to replay, use limits retention instead. The instructive test: run nats consumer next without --ack, and fetch again after 30 seconds — the same messages return, demonstrating redelivery. Check nats consumer info JOBS WORKER to see pending and redelivered counts.
Client libraries (Go, Java, Python, JavaScript, Rust, C#) expose the same concepts programmatically, but API names have shifted across releases — newer libraries favor a simplified consume/fetch style over older patterns. Confirm the current API against your library's release notes before copying any sample code, including this one.
The trade-offs you sign up for
Duplicates are your problem. At-least-once means a message can arrive twice: after an AckWait expiry, after a crash between processing and acking, or after a network blip that ate the ack. Handlers must be idempotent — deduplicate on a message ID (JetStream supports publisher-set dedup IDs via Nats-Msg-Id) or make the side effect naturally repeatable, like an upsert keyed on the job ID.
AckWait sizing is a real decision. Too short, and normal processing latency triggers spurious redeliveries — your workers do the same job twice and duplicates flood downstream. Too long, and recovery from genuine failures stalls. Size it to worst-case batch processing time plus a comfortable margin, and measure actual processing latency before tuning.
Pull adds a round trip. Fetching is a request/response cycle, so per-message latency is slightly higher than a push subscription. For high-throughput workers this is usually irrelevant — you amortize it with batches — but for latency-sensitive single-message dispatch, a push consumer with a queue group may fit better.
Misconfiguration silently discards. Stream storage limits and MaxDeliver caps will drop messages if they don't match your workload. Watch nats stream info for discarded-message indicators in any serious deployment.
Where to start
Spin up nats-server -js locally, run the CLI sequence above, and do the kill-mid-batch experiment: fetch without acking, wait out the AckWait, and watch the redelivery. Ten minutes of that teaches the semantics better than any diagram. Then pick one task-queue-shaped workload in your system — the one where a slow worker currently means lost or duplicated work — and prototype it behind a durable pull consumer with idempotent handlers. That's the shape pull consumers were built for.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.