The Recommended Strategy
To guarantee a stable full-dataset retrieval in the Confluence REST API, you must implement client-side deduplication using a Set or Hash Map of unique identifiers (e.g., pageId). Because the API uses offset-based pagination (the start parameter) and lacks cursor-based bookmarks, it cannot guarantee that items in page N + 1 are disjoint from page N if the underlying data changes during the request sequence.
Why Duplicates Occur
The API does not provide a snapshot of the dataset at the start of the pagination process. It performs a fresh query for every request. If a page is deleted or added at an index lower than your current start value, the entire index shifts:
- Insertions: A new item added at index 0 pushes an existing item from index 999 to 1,000. If you just finished fetching index 0–999, that item will appear again in your next request (1,000–1,999).
- Deletions: An item removed at index 0 pulls the item at index 1,000 into the 999 slot. You will skip that item entirely.
Implementation Steps for Stable Retrieval
- Initialize a Tracking Set: Create a collection to store the IDs of all processed entities.
- Sequential Fetching: Use a
while loop to increment the start parameter by the limit value (e.g., 0, 1000, 2000) until the API returns an empty list or the next link is absent.
- Filter on Arrival: For every item in the response, check if the ID exists in your tracking set. If it does, discard it; otherwise, add it to your dataset and the set.
- Handle Timeouts: For very large datasets, high
start values can increase server latency. If you encounter timeouts, reduce the limit (e.g., to 250) to lower the per-request load.
Verification Command
You can verify the behavior of the offset by running two sequential calls and comparing the last item of the first set with the first item of the second set:
# Request 1
curl "https://your-domain.atlassian.net/wiki/rest/api/content?start=0&limit=50"
# Request 2
curl "https://your-domain.atlassian.net/wiki/rest/api/content?start=50&limit=50"
Environmental Differences
In Confluence Cloud, the API is strictly managed by Atlassian's infrastructure, and the 1,000-result cap is a hard limit. In Confluence Server/Data Center, the maximum limit may be configurable via system properties in server.xml or the admin console, but the fundamental risk of offset-shifting remains identical.
Missing Diagnostic: Are you performing these retrievals during a known maintenance window or during peak user activity? If the dataset is highly volatile, client-side deduplication is the only viable path; if it is static, a simple sequential loop is sufficient.