Configure Harbor Cross-Region Replication for Disaster Recovery
Set up Harbor v2.8+ replication between two registries for cross-region disaster recovery: robot accounts, a target endpoint, a scheduled tag-filtered policy, digest checks and a promotion checklist.
16 Jun 2026, 19:00 UTC

What you are building
A Harbor instance in a second region that holds the same release images as the primary, so a DNS or load-balancer change is enough to fail over. Harbor's replication feature copies projects, repositories and tags to a remote registry endpoint, and from v2.7 onward it can also carry vulnerability scan results and SBOMs. This guide covers a scheduled push policy from primary to secondary on Harbor v2.8 and later, plus the checks that tell you the copy is actually usable.
Replication is not a backup of Harbor itself. It moves artifacts. Users, robot accounts, project configuration, quotas and webhook settings stay behind and must be managed separately.
Prerequisites
- Two Harbor instances, v2.8 or later, each reachable over HTTPS. The source must trust the target's certificate; if you use an internal CA, add it to the source's trust store rather than disabling verification.
- A robot account on the target with push and pull permission on the destination projects. A source-side robot is only needed if you drive the API; the UI uses your admin session.
- Set robot token expiry to never (0), or rotate tokens on a schedule you actually follow. An expired token surfaces as an authentication error in the job log, not as a warning beforehand.
- Enough storage on the target for the filtered subset plus headroom. Replication never prunes the target.
Step 1: Register the target registry
In the source Harbor UI open Administration > Registries > New Endpoint.
- Provider: Harbor
- Endpoint URL:
https://target-harbor.example.com - Access ID / Secret: the target robot account credentials
- Verify Remote Cert: leave enabled when the CA is trusted
Use Test Connection and expect a success response before saving. Name the endpoint something region-specific, such as dr-eu-west, because every policy pointing at that region reuses it.
Step 2: Create the replication policy
Go to Administration > Replication > New Policy. A push-mode policy reads from the local instance and writes to the endpoint you just created.
Scope and filters
- Source project: pick the projects you actually need at failover time, not every project. Smaller scope means shorter jobs and less target storage.
- Destination namespace: leave empty to keep project names identical, which keeps image references unchanged after promotion. A prefix such as
dr-is safer for testing but requires rewriting manifests at cutover. - Tag filter: a pattern like
release-*keeps mutable development tags out of the DR registry. Harbor v2.8 also supports label selectors on OCI artifacts and Helm charts; older releases filter on tags only.
Trigger and advanced options
- Trigger: scheduled, for example
0 2 * * *for a daily 02:00 UTC run. Event-based (on push) replication reacts faster but can produce bursts of jobs when a CI pipeline pushes many tags at once. - Override existing resources: enable it if the same tag can be re-pushed with different content. Without it, the target keeps the first copy and your digests silently diverge.
- Replicate deletion: off by default. Turning it on makes the target delete artifacts when the source does, which is destructive and one-way. Enable only with a documented reason.
- Job timeout: raise it if you replicate multi-gigabyte images; the default can expire mid-transfer on slow cross-region links.
Step 3: Run it once by hand
Trigger the policy manually and watch Replication > Jobs. A healthy run moves through pending, in progress and succeeded. Open the job to see per-artifact rows with repository, tag, digest and duration.
The error column is the useful part when something breaks: authentication failures point at the robot token, connection errors at certificates or firewalls, and timeout errors at link speed or job timeout.
Checks that prove the copy is usable
Compare digests, not tags
Tags are mutable; digests are not. From a workstation with skopeo installed and credentials for both registries, inspect the same reference on each side:
skopeo inspect docker://source-harbor.example.com/production/api-service:release-1.2.0 --creds 'robot$source:SOURCE_TOKEN' | jq -r .Digest
skopeo inspect docker://target-harbor.example.com/production/api-service:release-1.2.0 --creds 'robot$target:TARGET_TOKEN' | jq -r .DigestIdentical digests mean the artifact content matches. A mismatch usually means override was disabled and the target held an older manifest, or that something pushed to the target outside the policy.
Confirm the target can serve a pull
Point a test client at the target, pull the image and start it. This catches problems digest comparison misses, such as a target project the robot cannot read or a missing proxy cache configuration.
Check job history and metrics
Harbor exposes replication counters to Prometheus. A query such as harbor_replication_job_total{status="success"} should increase after each scheduled run; alert on failed counts over a rolling window rather than waiting for the next manual check. Confirm the exact metric names and labels against the metrics endpoint of your own build, since exported metric sets change between releases.
Recovery options
- Network or certificate failure: fix the underlying cause, then re-run the policy. Replication is idempotent, so a re-run does not duplicate artifacts.
- Target storage filling: run garbage collection on the target and review the policy scope. GC removes untagged blobs; it does not remove tags you chose to replicate.
- Wrong filter or namespace: edit the policy and re-run. Artifacts already copied under the old scope stay on the target and must be removed deliberately.
- Large-image timeouts: increase the job timeout, or split the policy so very large repositories replicate on their own schedule.
Promotion checklist
- Confirm every image your workloads reference exists on the target with a matching digest.
- Recreate the accounts, projects and configuration the target needs, since replication does not carry them.
- Repoint DNS or the load balancer at the target, and verify a pull through the normal hostname.
- Update CI/CD credentials to target robot accounts on the promoted instance.
- Decide how the two instances will be reconciled when the primary returns; without that plan you can end up with divergent tags in both directions.
Limitations to plan around
- Robot accounts, users, OIDC or LDAP settings, project quotas, retention rules and webhook configuration are not replicated.
- Scan results and SBOMs replicate only from v2.7 onward and require a scanner configured on both sides.
- Replication is asynchronous. The target is only as current as the last successful job, so the schedule defines your recovery point objective.
- Deletion replication is destructive and does not restore anything.
What to verify in your own environment
Exact menu labels, the availability of label-based filtering, which scan data is replicated, and the Prometheus metric names all vary by Harbor release. Check them against the documentation for the version you run before relying on this configuration, and treat the version-specific statements here as assumptions to confirm rather than guarantees.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.