tf.data pagination with .skip()/.take() – unresolved shuffle ordering behavior
26.5K reputation · 07 Dec 2025, 03:22 UTC
tf.data pagination with .skip()/.take()
Goal: Retrieve a fixed window from a large dataset for training or inference, e.g., page 3 of size 100 via .skip(200).take(100). This approach is documented in the tf.data API.
Unresolved behavior
When a shuffle operation precedes the pagination steps, the resulting page contains a random sample of the original data. If shuffle follows paging, the page is a contiguous slice of an already shuffled stream. The tf.data documentation does not specify which ordering preserves epoch‑level determinism or how the shuffle buffer is managed across epochs when combined with .skip() and .take(). Additionally, applying .take() to an unbounded source can result in an unknown cardinality, potentially breaking downstream operations that expect a fixed size.
Questions
- Does shuffling before pagination guarantee that the same page contents appear in every epoch, or does shuffling after pagination produce a deterministic page per epoch?
- What buffer size is required for .shuffle(buffer_size) to ensure that a paged window contains unique, non‑repeating elements across epochs?
- When .take() is applied to an unbounded source, how can one reliably determine the resulting dataset’s cardinality for subsequent batch or repeat operations?