Setting Up Geo-Replication in Apache Pulsar for Multi‑Region Low‑Latency Messaging
Learn how to configure Apache Pulsar’s geo‑replication to copy messages from a source cluster to multiple regional targets, preserving order and enabling local low‑latency consumption.
05 Apr 2026, 12:13 UTC

Problem: latency and single‑point‑of‑failure across data centers
When an application spans multiple regions, a single Pulsar cluster forces all producers and consumers to traverse long‑distance network paths. This adds latency and creates a bottleneck: if the cluster goes down, every region loses messaging capability.
How Pulsar geo‑replication solves it
Pulsar’s built‑in replication subsystem copies messages from a source cluster to one or more target clusters automatically. Replication preserves per‑key ordering and provides eventual consistency, allowing consumers in each region to read from a local copy with sub‑second lag under normal conditions.
Worked example: replicating an “orders” topic
Assume three clusters: us-east (source), us-west and eu-central (targets). A producer in us-east publishes to the topic persistent://public/default/orders using a message key such as the order ID.
1. Register the clusters
# Run on any broker host with admin privileges
pulsar-admin clusters create us-east \
--broker-url pulsar://us-east-broker:6650
pulsar-admin clusters create us-west \
--broker-url pulsar://us-west-broker:6650
pulsar-admin clusters create eu-central \
--broker-url pulsar://eu-central-broker:6650
These commands create cluster metadata in the source cluster’s configuration store. They require the superuser role or a role with admin permissions on the cluster metadata.
2. Enable replication for the namespace
pulsar-admin namespaces set-replication-clusters public/default \
--clusters us-east,us-west,eu-central
After this command, the source cluster begins to stream messages to the two target clusters. No broker restart is needed; the change is applied dynamically.
3. Produce and consume locally
In us-east:
# Producer (Java pseudo‑code)
producer.newMessage()
.key(orderId) // preserves ordering per key
.value(orderJson)
.send();
Consumers in us-west and eu-central subscribe to the same fully‑qualified topic name and receive messages with only the replication lag as delay.
Verification steps
- Check replication lag: In Pulsar Manager navigate to the “Replication” dashboard for the namespace
public/default. The “Replication Lag” metric for each target cluster should stay below a chosen threshold (e.g., 500 ms) under steady load. - Validate ordering: Produce a sequence of 1000 messages with keys
orderId=0..999. Consume from each target cluster and confirm that the received sequence matches the original order with no gaps. - Test resilience to network partition: Using a traffic‑shaping tool (e.g.,
tcon Linux), introduce latency or drop packets betweenus-eastandus-westfor a few minutes. Observe that replication lag rises and then returns to baseline after the partition heals, with no missing messages when consumption resumes.
All verification commands assume you have read‑only access to the monitoring endpoints and admin rights to run pulsar-admin.
Trade‑offs and limitations
- Storage overhead: Each replicated message persists in every target cluster, roughly multiplying storage usage by the number of replicas.
- Replication lag during partitions: If inter‑cluster bandwidth drops, lag can increase proportionally; extreme partitions may cause back‑pressure on the source.
- Schema compatibility: Schemas must be evolvable in the same way across clusters; a breaking schema change in the source will propagate and may cause deserialization errors on targets if they have not been updated.
- No exactly‑once guarantee across clusters: Retries at the producer level or replication restarts can lead to duplicate messages appearing in a target.
Actionable closing
Geo‑replication is a practical way to achieve low‑latency, fault‑tolerant messaging across regions when you can tolerate the extra storage and occasional lag. Start by registering your clusters, set the replication list on the namespace, and monitor lag via the Pulsar Manager dashboard. Verify ordering and resilience with the steps above before promoting the setup to production.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.