Diagnosing and Fixing Harbor Image Replication Failures
When Harbor image replication stalls or fails, pinpointing the root cause quickly saves time. This guide walks through common conditions—authentication errors, network timeouts, resource limits, and mis‑configured endpoints—provides ordered checks, actionable fixes, and escalation thresholds.
10 Jul 2026, 08:39 UTC

Why Replication Fails in Harbor
Harbor’s replication engine copies images from a source project to a target project or external registry. A failure can surface as a job stuck in pending, marked failed, or with a specific error code. The most common culprits are authentication problems (401/403), network reachability, pod resource constraints, and mis‑configured target endpoints.
Diagnostic Overview
Below is a quick reference table that maps a recognizable condition to its likely cause, the checks you should run, the fix, and when to seek escalation.
| Condition | Likely Cause | Checks | Fix | Escalation |
|---|---|---|---|---|
Job status = pending or failed with no progress | Authentication failure (401/403) | Verify credentials in Harbor UI, check replication.yaml, run harborctl replication list | Update token/username/password, rotate secrets, re‑create replication rule | After 2 retries, contact registry admin |
Job logs show timeout or connection refused | Network connectivity | Run curl -v https://TARGET_REGISTRY/v2/ from the Harbor host, check kubectl logs of replication pods | Fix network routes, open firewall ports (443), add DNS entries | If external registry unreachable, involve network ops |
Job stops with OOMKilled or CrashLoopBackOff | Resource limits | Check pod status: kubectl describe pod, inspect resource usage via kubectl top pod | Increase CPU/memory limits in replication.yaml or deployment, or scale replicas | After adjustment, monitor for stability; if persistent, review cluster capacity |
Job logs contain 404 or endpoint not found | Target endpoint mis‑configured | Validate target URL in replication rule, test via curl from Harbor node, confirm TLS certificates | Correct URL, update TLS secret, restart replication pods | If external registry requires custom CA, involve security team |
Step‑by‑Step Diagnostic Flow
- Open the Harbor UI
- Navigate to Projects → Replication and locate the job.
- Check the
Statuscolumn; note any error messages displayed. - Click the job ID to view detailed logs.
- Verify Replication Configuration
- Locate the replication rule in
/etc/harbor/replication.yaml(or via the UI). - Ensure the
sourceandtargetURLs are correct, including protocol (https://). - Check credential fields:
username,password, ortoken. Avoid hard‑coded secrets in logs. - Run
harborctl replication list(if harborctl is installed) to confirm the rule is active.
- Locate the replication rule in
- Test Network Connectivity from the Harbor Host
- SSH into the node running Harbor (or use
kubectl execinto theharbor-corepod). - Execute:
curl -v https://TARGET_REGISTRY/v2/ - Expected output: HTTP 200 OK. If you see a timeout or connection error, the registry is unreachable.
- Check DNS resolution:
nslookup TARGET_REGISTRY.
- SSH into the node running Harbor (or use
- Inspect Replication Pod Logs and Status
- Run:
kubectl logs -n harbor harbor-replication-pod-name - Look for stack traces,
401/403messages, ortimeoutwarnings. - Check pod events:
to see if it was OOMKilled or crashed.kubectl describe pod -n harbor harbor-replication-pod-name
- Run:
- Validate Resource Allocation
- Run:
kubectl top pod -n harbor harbor-replication-pod-name - Compare CPU/memory usage against the pod’s limits defined in
replication.yamlor the Helm chart. - If usage approaches limits, consider raising them or adding replicas.
- Run:
- Check TLS and CA Certificates
- For HTTPS targets, ensure Harbor’s node trusts the registry’s CA. If using a self‑signed cert, add it to
/etc/ssl/certs/ca-certificates.crtor mount a customca-bundlesecret. - Test with
curl --cacert /path/to/ca.crt https://TARGET_REGISTRY/v2/to confirm the handshake succeeds.
- For HTTPS targets, ensure Harbor’s node trusts the registry’s CA. If using a self‑signed cert, add it to
- Apply Fixes and Re‑run the Job
- After making configuration changes, restart the replication pod:
kubectl rollout restart deployment -n harbor harbor-replication - Re‑trigger the replication job from the UI or via
harborctl replication trigger. - Verify job completes successfully within a reasonable time frame.
- After making configuration changes, restart the replication pod:
Practical Example: Updating Credentials
Suppose a job fails with a 401 error. The replication rule in replication.yaml looks like this:
replication_rules:
- name: "copy-to-prod"
source: "https://harbor.example.com/api/v2.0"
target: "https://prod-registry.internal:5000"
credentials:
username: "old_user"
password: "old_pass"
Replace the credentials by editing the file:
sed -i 's/old_user/new_user/' /etc/harbor/replication.yaml
sed -i 's/old_pass/new_pass/' /etc/harbor/replication.yaml
After editing, back up the file:
cp /etc/harbor/replication.yaml /etc/harbor/replication.yaml.bak
Restart the replication deployment:
kubectl rollout restart deployment -n harbor harbor-replication
Trigger a new job and observe it finish. If it still fails, verify the new credentials are valid by logging into the target registry directly from the Harbor host.
Escalation Criteria
- After two consecutive failed attempts with the same error, involve the registry administrator.
- If network diagnostics reveal firewall or routing issues, contact the network operations team.
- Persistent OOMKilled pods after raising limits should prompt a cluster capacity review.
- Mis‑configured TLS that requires custom CAs should involve the security or compliance teams.
Limitations and Verification
Harbor’s replication logs may not surface all underlying issues; external factors like upstream registry throttling can also cause timeouts. Always confirm changes by re‑triggering the job and checking both Harbor UI and kubectl logs. If the problem recurs, consider enabling debug logging in Harbor’s harbor-core container for deeper insight. Remember to keep a backup of replication.yaml before editing, and avoid exposing credentials in command history or unsecured files.
Conclusion
By following the ordered checks—UI inspection, config validation, network testing, pod log analysis, and resource verification—you can efficiently isolate and resolve most Harbor replication failures. Apply the fixes, re‑run the job, and if the issue persists beyond the defined escalation thresholds, bring the appropriate team into the loop.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.