Jira Data Center Index Replication Latency During High-Concurrency Writes
0 reputation · 06 Apr 2023, 03:42 UTC
Context
A Jira Data Center cluster uses a shared home directory with index replication enabled to keep search indexes consistent across nodes. Background indexing is active to avoid blocking user operations during re-indexing.
Goal
Achieve predictable search result parity across all cluster nodes immediately after a high-volume batch of issue updates, such as a bulk transition or import operation.
Constraints and Uncertainty
The documentation describes index replication as asynchronous but does not specify the maximum replication lag under sustained write throughput. JVM heap sizing for the Lucene index writer is configured per node, yet the interaction between local flush intervals and the replication protocol during concurrent writes remains undefined. It is unclear whether the replication mechanism guarantees eventual consistency within a bounded time window or if certain write patterns can cause indefinite divergence.
Questions
- What is the documented or observed upper bound for index replication latency across nodes when the cluster processes more than 500 issue updates per second?
- Does the replication protocol include a back-pressure signal that pauses local indexing when a replica falls behind, or can the primary node continue accepting writes indefinitely?
- Are there configurable thresholds (e.g.,
index.replication.max.lag) that control failover behavior when replication lag exceeds a defined limit?