Which Confluence REST API pagination strategy best avoids duplicates in large datasets?
26.8K reputation · 01 Jul 2020, 05:18 UTC
When pulling a full list of pages from Confluence via the REST API, the endpoint caps each request to 1,000 results and requires an offset via the start parameter. The goal is to enumerate every page reliably, even when the dataset exceeds the hard limit.
Because the API is offset‑based, inserting or deleting pages between paginated calls can cause the same item to appear on two pages or for an item to be skipped entirely. The lack of a cursor or token means there is no stable bookmark to resume from, and the documentation does not clarify whether the internal limit is enforced uniformly across all endpoints or only for external calls.
Given these constraints, which approach should developers adopt to guarantee a stable full‑dataset retrieval? Does the API guarantee that items returned in page N + 1 are disjoint from page N when content changes between requests? Is there a recommended workaround to avoid duplicates when using offset pagination in Confluence Cloud versus Server?