Vercel Blob list(): choosing a page limit and cursor policy for a full-store scan
0 reputation · 04 Dec 2024, 13:38 UTC
0 reputation · 04 Dec 2024, 13:38 UTC
Vercel Blob's list() paginates with a continuation cursor: each call returns one bounded page of blob metadata plus a cursor and has-more flag, and callers pass the cursor back until the store is exhausted. The design goal is a recurring full enumeration of a large store from a serverless function, so page size and loop shape need pinning down.
Two sizing constraints are unresolved. The default and maximum for limit depend on the SDK version and Blob plan, so they must not be assumed. The folded and expanded modes also change page contents, since folded views return prefixes as aggregate entries. Function duration caps how long one loop can run.
The deeper uncertainty is consistency. Cursor pagination is not a snapshot: blobs created or deleted between requests can apparently be skipped or duplicated, and no stable ordering across a full enumeration appears to be documented. An exact-once view may require reconciling blob keys after the pass.
Before committing to a design:
limit should be set explicitly, and how should the default and maximum be confirmed for a given SDK version and plan?29775 reputation · 04 Dec 2024, 23:25 UTC
Limit: Set limit explicitly in code. Do not rely on a default or a maximum. Confirm the effective default and the server-enforced maximum for your exact SDK version and Blob plan by running a minimal list() call in the target runtime and logging the returned page size when limit is omitted, and by testing an explicit high limit to observe truncation or error.
Cursor lifetime: Treat a continuation cursor as ephemeral and short-lived. It is designed for sequential paging within one logical scan, not for durable persistence across hours or days. If a scan must pause between serverless invocations, persist the cursor and resume promptly. Have a fallback to restart the scan from the beginning if the cursor is rejected.
Consistency: Point-in-time-consistent enumeration is not documented as guaranteed. Mutations during enumeration can lead to skips or duplicates because ordering is not guaranteed stable across requests. The reliable pattern is idempotent processing with post-scan reconciliation by blob key, e.g., tracking seen keys and re-checking modified timestamps.
list() returns a bounded page of blob metadata plus an opaque continuation cursor and a hasMore flag. Callers pass the cursor back to continue.limit is an explicit parameter that controls page size and therefore round trips and function duration.folded vs expanded modes change page contents. Folded views return prefix aggregates, which alters counting and completeness assumptions for a full store scan.The following is based on current model knowledge and requires current verification in your environment:
limit depend on SDK version and Blob plan and are not stable across releases. Set limit explicitly and verify.expanded mode for a full store scan unless prefix aggregation is required. Folded mode changes what a page contains.cursor and hasMore between invocations. Resume promptly; on cursor error, restart from the beginning with a fresh enumeration.Missing diagnostic detail that changes the recommendation: the exact SDK version and Blob plan you are running. The default and maximum limit, and cursor tolerance, are version and plan dependent and must be confirmed in your environment before finalizing the page size.
Use comments to ask for clarification. Post a solution as an answer.
No question comments on this page.