RediSearch LIMIT clause: consistent pagination across concurrent modifications?
0 reputation · 06 Oct 2023, 12:42 UTC
0 reputation · 06 Oct 2023, 12:42 UTC
We want to fetch a specific page of search results from a large RediSearch index while ensuring that the page contents remain consistent even if documents are added or removed during pagination.
RediSearch supports a LIMIT offset,count clause, but the index is mutable; new documents can be inserted or deleted between successive queries. The SCAN family offers cursor‑based iteration over key collections, yet it provides no guarantees on ordering or non‑overlap when the underlying set changes.
Documentation states that a cursor is a snapshot of the data at the time of the first SCAN call, but it does not explicitly state whether that snapshot is used for subsequent iterations or whether the order is stable. Likewise, RediSearch’s LIMIT does not mention handling of concurrent modifications.
LIMIT offset,count will return the same set of documents if the index is modified between paginated requests?SCAN (or its variants) for pagination, is the cursor guaranteed to reflect a consistent snapshot of the set, preventing duplicates or gaps across pages when the set changes?29775 reputation · 06 Oct 2023, 13:17 UTC
1. LIMIT offset,count does not provide a stable snapshot. The query sees the index as it exists at the moment the command starts; inserts or deletes that happen before the command finishes can shift which documents fall into the requested offset.
2. SCAN cursors capture the key set at the first call, but the cursor does not guarantee that the order remains stable or that no keys are duplicated or omitted if the set changes between iterations.
3. RediSearch exposes no built‑in feature that guarantees deterministic pagination over a mutable index. Typical patterns involve a stable sort key, client‑side offset recalculation, or routing pagination traffic to a read‑only replica.
RediSearch’s query engine is essentially “read‑as‑you‑go”: it calculates the result set on the fly using the current state of the index. Because the index is mutable, the offset used in LIMIT is applied to whatever set of documents is present when the query starts. If a new document that would appear before the offset is inserted, the first page will now contain that document, and the previously first document will shift to the second page. The same applies to deletions.
The SCAN family is designed for iterating over keys, not for stable pagination. The cursor returned by SCAN is a simple token that the server uses to resume scanning; the server does not freeze the key space. Therefore, changes to the key set can lead to duplicates or gaps across pages.
LIMIT into a snapshot‑isolated operation.LIMIT to the desired page size. This is a cursor‑style pagination that is immune to inserts before the lastKey.RediSearch 1.x behaved similarly, but newer releases may introduce optional snapshot isolation for LIMIT. Please confirm the exact module version you are running (e.g., 2.0.x, 2.1.x, or a custom build). That information will determine whether any new flags or configuration options exist that could change the behavior described above.
# Create a small index
FT.CREATE idx ON HASH PREFIX 1 doc: SCHEMA title TEXT
# Insert documents 1-5
HSET doc:1 title "A"
HSET doc:2 title "B"
HSET doc:3 title "C"
HSET doc:4 title "D"
HSET doc:5 title "E"
# Fetch page 1 (offset 0, count 2)
FT.SEARCH idx "*" LIMIT 0 2
# Insert a document that would sort before the current first page
HSET doc:0 title "Z"
# Fetch page 1 again
FT.SEARCH idx "*" LIMIT 0 2
Observe whether the first page now includes the newly inserted document.
RediSearch’s LIMIT and SCAN cursors are not designed for deterministic pagination over a mutable dataset. Use a stable sort key with client‑side offset tracking, or route pagination to a read‑only replica, to achieve consistent results across page boundaries.
Use comments to ask for clarification. Post a solution as an answer.
No question comments on this page.