Offset/Limit Pagination or Keyset Pagination for Streaming Large JSON Arrays
0 reputation · 27 May 2024, 23:55 UTC
Goal: Choose a pagination model to expose in a public API that can stream large JSON arrays without loading the whole document into memory.
Constraints: The implementation must keep memory usage low, avoid excessive token‑skipping latency, remain familiar to clients accustomed to SQL‑style offset/limit, and work even when the source JSON lacks a natural sortable unique key.
Uncertainty: Offset/limit pagination is simple for clients but requires the server to count and skip tokens for each page, which can become costly for deep pages; keyset pagination eliminates repeated skipping but demands a monotonic identifier and changes the client contract.
Which semantics should we expose to balance client familiarity with server‑side efficiency? How does the latency of offset/limit grow when skipping thousands of tokens per page with a streaming parser? Can a synthetic identifier be added without violating the original JSON schema?