Diagnosing and Fixing Index Replication Lag in Jira Data Center
When Jira search becomes stale or dashboards lag, a hidden index replication issue may be at play. This guide walks you through symptoms, root causes, verification steps, targeted fixes, and when to raise an escalation.
27 Feb 2026, 15:43 UTC

Recognizable Symptoms
When the search engine or dashboards show stale data, the following signs often point to an index replication problem:
- Search results or JQL queries omit recently updated issues.
- Dashboard gadgets display counts that lag behind the actual issue count.
- Logs contain
IndexReplicationQueuebacklog warnings orLucene index corruptionstack traces. - HTTP 500 errors from the search service rise above 1 % for more than 10 minutes.
Root‑Cause Mapping
| Cause | Typical Symptom | Diagnostic Indicator |
|---|---|---|
| Network partition between nodes | Stale search across the cluster | Replication queue grows on all nodes |
| Shared‑home snapshot stall | Missing recent updates | SnapshotManager logs show timeout or retry loops |
| JVM GC pause >5 s | Search latency spikes | GC logs show long pauses around snapshot times |
| Disk I/O saturation on shared home | Replication queue backlog | High I/O wait in iostat or vmstat |
| Version mismatch between nodes | Index corruption on one or more nodes | Logs show incompatible index format errors |
| Manual rebuild without coordination | Split‑brain state, inconsistent search | Different replicationIndexVersion per node |
Verification Checklist
- Replication status on each node
Run on every node (replace<JIRA_URL>and<API_TOKEN>):
Check thatcurl -u <API_TOKEN> -X GET <JIRA_URL>/rest/api/2/cluster/index/replication/statusreplicationIndexVersionmatches across nodes andqueueSizeis zero. - Log inspection
Searchatlassian-jira.logforIndexReplicationManagerandSnapshotManagerentries. Correlate timestamps with GC logs (e.g.,gc.logorjvm.log). - Shared‑home health
Verify mount options and usage:
Ensure the file system is block‑based (EBS, Persistent Disk) and not NFS with burst latency.mount | grep shared-home df -h /var/atlassian/application-data/jira/shared-home - Version consistency
Confirm all nodes run the same Jira major/minor version and Java runtime. On each node:./bin/atlassian-jira.sh -v java -version - Integrity check
Optionally run a full integrity scan (may be resource intensive):
Expectcurl -u <API_TOKEN> -X POST <JIRA_URL>/rest/api/2/cluster/index/replication/integrityOKon all nodes.
Common Fixes
- Restore network connectivity
Verify that all nodes can reach each other on the replication port (default 8080). Usepingornc -zvfrom each node. - Force a snapshot
If the queue is stuck, trigger a manual snapshot (available on JDC 8.20+):
Monitor logs for successful snapshot completion.curl -u <API_TOKEN> -X POST <JIRA_URL>/rest/api/2/cluster/index/replication/snapshot - Move snapshots to local SSD
Setjira.index.replication.local=trueinjira-application.propertieson each node, then restart. This keeps snapshots on local storage, reducing shared‑home I/O. - Tune GC and heap
Increase heap size or switch to G1GC if pauses exceed 5 s. Apply changes injira-application.propertiesand perform a rolling restart. - Provision dedicated IOPS for shared home
Ifiostatshows high wait times, allocate an SSD with sufficient IOPS or move the shared home to a dedicated block device. - Coordinated rebuild
When rebuilding the index, pause all follower nodes, rebuild on the leader, then restart followers one by one. Avoid deleting$SHARED_HOME/indexwithout coordination.
Escalation Thresholds
If after 30 minutes the replication queue remains above 500 k ops, or if integrity checks report CORRUPT on more than one node, elevate the issue to Atlassian support. Also raise an alert if search‑related HTTP 500 errors persist above 1 % for more than 10 minutes.
Post‑Fix Validation
- Re‑run the replication status API; all nodes should report identical
replicationIndexVersionandqueueSize=0. - Execute a representative JQL query (e.g.,
updated >= -1h) on each node via REST and compare result counts. - Monitor
atlassian-jira.logfor 15 minutes; no newIndexReplicationQueueWARN entries should appear. - Run the integrity API again; all nodes should return
OK. - Verify dashboard gadgets update within 30 seconds of an issue transition.
Limitations and Practical Checks
- Manual snapshot API is only available on Jira Data Center 8.20+. Earlier releases require a node restart to force a snapshot.
- Enabling local index replication increases per‑node disk usage by roughly the size of the index; plan capacity accordingly.
- Deleting
$SHARED_HOME/indexforces a full rebuild and can take hours on large installations (≈ 1 M issues per hour on SSD). Schedule during low‑traffic windows. - GC tuning changes require a rolling restart and should be tested in a staging environment that mirrors production data volume.
- Integrity scans consume CPU and memory; run them during off‑peak periods.
By following this ordered diagnostic workflow, you can isolate the root cause of index replication lag, apply the appropriate remediation, and verify that the cluster returns to a healthy state before resuming normal operations.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.