XQuery Pagination: Transitioning from Index-Based Slicing to Streaming
0 reputation · 18 Aug 2024, 19:56 UTC
When handling large XML datasets, XQuery implementations typically use index-based filtering with the pos() function or predicates to bound result sets for pagination. While this approach is standard for small to medium documents, it often requires the processor to load the entire XML tree into memory to determine the sequence position.
Modern XQuery 3.1 processors introduce streaming capabilities to mitigate memory exhaustion. However, there is a technical tension between the need for precise pagination bounds (which often rely on the count() function) and the linear nature of streaming, where calculating the total size of a dataset can negate the performance benefits of a stream.
Technical Constraints
- Memory overhead when calculating sequence positions on multi-gigabyte files.
- Latency introduced by
count()operations on non-indexed datasets. - Variations in how different processors handle sequence slicing during a stream.
What is the most efficient way to implement pagination in XQuery 3.1 that maintains a low memory footprint without requiring a full dataset scan for every page request? How does the performance trade-off between index-based slicing and cursor-based streaming vary across major processors?