Enable Asynchronous Replication in Norg for High Availability
Learn how to enable Norg’s asynchronous replication, set a 3‑node factor, monitor lag, and recover from node outages with step‑by‑step commands and verification checks.
29 Apr 2026, 19:25 UTC

Desired Outcome
Turn on Norg’s built‑in asynchronous replication, set a replication factor of 3, and verify that writes are propagated quickly while the cluster remains available if a node fails.
Prerequisites
- Running Norg server version 3.2 or newer on all nodes.
- Cluster nodes reachable over a secure network (TLS + authentication).
- Sufficient disk space and CPU headroom for the replication thread.
- Prometheus or another metrics scraper configured to hit the
/metricsendpoint on each node.
Enabling Asynchronous Replication
Replication is controlled via the norg.conf file and the norgctl CLI. First, edit the configuration on every node.
# /etc/norg/norg.conf
replication.enabled = true
# Optional: limit the replication thread to 2 CPUs per node
replication.thread_cpu_limit = 2
Then, on each node run the following command as the Norg system user:
# norgctl replication enable --factor 3 --tls --auth
# --factor 3 sets the replication factor.
# --tls ensures traffic is encrypted.
# --auth enables node‑to‑node authentication.
After the command completes, restart the Norg process to apply changes:
# systemctl restart norg
Configuring the Replication Factor
The replication factor determines how many replicas each key is stored on. A factor of 3 is recommended for production to tolerate two simultaneous node failures.
# norgctl replication set-factor 3
Verify the setting with:
# norgctl replication status
Replication factor: 3
Mode: asynchronous
Securing Replication Traffic
Replication traffic must be protected. Ensure each node’s tls.enabled is true and that a shared TLS certificate bundle is present. Authentication can be enabled via a shared key file:
# /etc/norg/auth.key
Configure the key path in norg.conf:
auth.key_file = "/etc/norg/auth.key"
Verifying Replication Health
After enabling replication, monitor the replication_lag_seconds metric. A good target is <1 second for most workloads.
# Prometheus query
avg_over_time(replication_lag_seconds{cluster="prod"}[5m])
Run a write‑heavy workload (e.g., 10,000 inserts per second) and watch the lag converge. Example check script:
#!/usr/bin/env bash
THRESHOLD=1
while true; do
LAG=$(curl -s http://node1:9100/metrics | grep replication_lag_seconds | awk '{print $2}')
if (( $(echo "$LAG < $THRESHOLD" | bc -l) )); then
echo "Replication lag $LAG s – OK"
break
else
echo "Lag $LAG s – waiting…"
sleep 10
fi
done
Handling Node Failures
Simulate a node outage by stopping the Norg process on one replica:
# systemctl stop norg
Read traffic should continue to be served by the remaining nodes. Writes will be queued in the replication backlog until the failed node recovers. Verify with the replication_backlog_size metric:
# Prometheus query
replication_backlog_size{node="node2"}
When the node comes back online (systemctl start norg), the replication thread will catch up automatically.
Read‑Repair Strategy
Because Norg uses eventual consistency, you can trigger a read‑repair on a key that has diverged:
# norgctl read-repair key123
Repair status: success – majority value restored.
For bulk repair, schedule a nightly job that scans for keys with replication_lag_seconds > 5 and runs read-repair on them.
Monitoring and Alerting
Expose the following metrics on each node:
| Metric | Description |
|---|---|
| replication_lag_seconds | Time since the last write was replicated to this node. |
| replication_backlog_size | Number of pending replication events. |
| node_cpu_usage_percent | CPU load of the replication thread. |
Set Prometheus alerts:
ALERT ReplicationLagHigh
IF avg_over_time(replication_lag_seconds{cluster="prod"}[5m]) > 1
FOR 5m
LABELS {severity="critical"}
ANNOTATIONS {summary="Replication lag > 1s"}
Known Limitations
- Eventual consistency: reads may return stale data until read‑repair runs.
- High write throughput can spike CPU usage on the replication thread; monitor
node_cpu_usage_percent. - Replication traffic is only protected if TLS and authentication are correctly configured.
Recovery and Rollback
If replication causes performance regressions, you can disable it quickly:
# norgctl replication disable
# Restart Norg
# systemctl restart norg
To roll back to a previous configuration, restore the old norg.conf from backup and restart the service.
Diagram
Illustration of a 3‑node asynchronous replication cluster.
Node A → Node B
Node A → Node C
Node B ↔ Node C (replication sync)
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.