Filtered Replication vs View-Based Extraction: Which Minimizes CPU Load During Large Syncs?
0 reputation · 02 Jun 2020, 02:18 UTC
When synchronizing large CouchDB datasets, the choice between filtered replication and view-based extraction can significantly affect resource consumption. The goal is to keep CPU usage on the source node within acceptable limits while still delivering only the desired subset of documents to the target.
Filtered replication evaluates a JavaScript filter for every document considered for replication, which can become expensive as the database grows. View-based extraction, on the other hand, relies on pre-built design documents and indexes, potentially reducing per-document overhead but requiring an additional query step before initiating the replication.
Uncertainty remains about which approach yields lower overall CPU load and faster completion times for typical large-scale syncs. Additionally, it is unclear how replication latency scales when the filter logic is complex versus when the view index is highly selective.
How does the per-document sandbox execution in filtered replication impact CPU usage compared to the pre-filtering cost of view-based extraction? What are the trade-offs in replication latency and resource consumption between the two methods? Can view-based extraction fully replace filtered replication for selective data movement without increasing round‑trip overhead?