Designing a Minimal Couchbase XDCR Deployment Across Two Data Centers
Outlines the requirements, smallest viable design, trust boundaries, operational checks, failure modes, and triggers for redesigning a Couchbase XDCR link between data centers.
22 Sept 2025, 01:33 UTC

Requirements
To run Couchbase Cross Datacenter Replication (XDCR) you need:
- Two or more Couchbase clusters, each running version 6.5 or newer (7.0+ recommended).
- Network latency between data centers consistently under 100 ms for synchronous conflict resolution.
- Compatible bucket names and, if using scopes/collections, the same scope names on source and target.
Smallest suitable design
A single bidirectional XDCR pair linking one bucket in each direction is enough to test active‑active or active‑passive setups.
- Create a source bucket srcBucket on cluster A and a target bucket tgtBucket on cluster B.
- Enable XDCR from A → B for srcBucket → tgtBucket using the default timestamp‑based conflict resolution.
- Repeat the step in the opposite direction (B → A) to obtain bidirectional flow.
- Limit replication to only the required scopes/collections (for example scope1.collection1) to reduce metadata overhead.
Example REST call (run on cluster A, replace placeholders with actual values):
curl -u Administrator:password -X POST http://nodeA:8091/pools/default/buckets/srcBucket/xdcr/replicate -H 'Content-Type: application/json' -d '{
"fromBucket":"srcBucket",
"toClusterReference":"clusterB-ref",
"toBucket":"tgtBucket",
"replicationType":"continuous",
"conflictResolution":"timestamp"
}'
The same call with source and target swapped creates the reverse link.
Trust/data boundaries
Treat each data center as an independent trust zone.
- XDCR traffic is encrypted with TLS; the source cluster initiates the TLS handshake using certificates that must be valid on both sides.
- Access control is enforced by RBAC on the source cluster: only users with the xdcr_admin role can create, modify, or delete replication definitions.
- The destination cluster applies the same bucket‑level privileges that exist locally; it does not inherit source permissions.
Operational checks
Regularly verify the health of the replication link:
- Replication lag – via UI → XDCR → Replication tab or the REST endpoint /pools/default/replications. Look for the field replication lag and ensure it stays below your SLA (for example <5 seconds) during peak load.
- Checkpoint persistence – confirm that the checkpoint value advances in the XDCR logs; a stalled checkpoint indicates a problem.
- TLS validity – grep "SSL handshake completed" /opt/couchbase/var/lib/couchbase/logs/xdcr.log on both clusters and ensure the output contains a timestamp.
- Mutation error rate – monitor system eventing logs for entries like XDCR conflict resolution failed or Mutation rejected. A rising rate may signal clock skew or quota exhaustion.
Failure modes
Common ways the XDCR link can degrade or stop:
- Network partition – if the inter‑DC path drops, the source continues to accept writes but replication pauses; writes accumulate in the source’s mutation queue until connectivity returns.
- Divergent system clocks – timestamp‑based conflict resolution relies on wall‑clock time; a skew of more than a few milliseconds can cause the wrong version to win, leading to unexpected data loss.
- Target bucket quota exhaustion – when the destination runs out of RAM or disk quota, XDCR throttles or rejects incoming mutations, creating back‑pressure on the source.
Design change conditions
You may need to revisit the architecture when any of the following occurs:
- A newer Couchbase release offers advanced conflict resolution (for example merge functions) that better fits your data model.
- The workload shifts to active‑active across more than two sites, requiring a full‑mesh or hub‑and‑spoke XDCR topology.
- Your tolerance for data loss tightens (for example financial transactions) and you consider switching to synchronous replication mode or adding a third‑party commit log.
- Operational overhead from monitoring many scopes/collections becomes prohibitive; you consolidate to a smaller set of replicated objects.
Limitations and practical verification
Remember that XDCR does not guarantee zero loss; during a network partition the source may continue accepting writes that are not replicated until the link heals, creating a potential data‑loss window.
To confirm the design is working as intended:
- Check the UI: XDCR status should show Running for both directions and replication lag <5 s under load.
- Run a quick API query: GET http://nodeA:8091/pools/default/replications and verify each entry has "status":"running" and a reasonable replication lag value.
- Validate TLS: grep "SSL handshake completed" /opt/couchbase/var/lib/couchbase/logs/xdcr.log on both clusters and ensure the output contains a timestamp.
- Audit RBAC: curl -u Administrator:password http://nodeA:8091/settings/rbac/users/local and confirm only intended users have the xdcr_admin role.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.