Direct Answer
For multi-gigabyte datasets in Wolfram Language, range-constrained Import generally uses less peak kernel memory than API-driven streaming when the source file is local or on a fast network mount, because it reads only the requested row window into a single in-memory expression. API streaming (URLRead/URLExecute with chunked processing) shifts memory pressure to the HTTP layer but introduces deserialization overhead per chunk and cannot match the columnar type inference of a single Import["Data", {start, end}] call.
Memory Overhead Trade-offs
Import with Range Constraints
- Peak memory: roughly the size of the requested row window × column count × Wolfram expression overhead (each number ≈ 16–24 bytes, each string ≈ 50+ bytes plus characters).
- No streaming buffer: the kernel allocates the full result expression at once; no intermediate chunks are retained.
- File handle: held only during the read; suitable for local SSD or high-throughput NFS.
- Limitation:
{start, end} works reliably for "CSV", "TSV", "Table"; for "JSON" or "XML" the entire structure is parsed first, negating the benefit.
API-Driven Streaming (e.g., URLRead + ImportString per chunk)
- Peak memory: chunk size × deserialization overhead + HTTP receive buffer (default 64 KB, configurable via
"ChunkSize").
- Sustained pressure: each chunk creates temporary expressions; garbage collection latency can cause transient spikes.
- Network dependency: latency, retries, and backpressure handling add complexity; rate limits may force smaller chunks, increasing per-row overhead.
- Advantage: works when the dataset is only accessible via HTTP API and cannot be staged locally.
Data Type Consistency with Mixed Columns
Range-constrained Import["Data", {start, end}] preserves per-column type consistency better. Wolfram's CSV/TSV importer scans the requested rows once, infers a single type per column (Integer, Real, String, DateObject, etc.), and coerces outliers to String. The same inference runs whether you import 10 rows or 10 million.
API streaming with per-chunk ImportString runs type inference independently on each chunk. A column that is all integers in chunk 1 but contains a single string in chunk 5 will produce {1, 2, 3} for the first chunk and {"1", "2", "3", "hello"} for the fifth — breaking downstream Dataset or numeric operations unless you manually unify types afterward.
Recommendation for Dynamic Remote Datasets
- If you can stage the remote data to a local file first (e.g.,
URLDownload to a temporary file, then range-import), do so. You get consistent typing, lower peak memory, and retryable reads.
- If staging is impossible (true streaming source, append-only log), use
URLRead[url, "Body" -> "Chunked"] with a fixed chunk size (e.g., 100k rows), parse each chunk with ImportString[#, "CSV"], and immediately Query/reduce to a compact representation (associations, packed arrays) before the next chunk arrives. Explicitly specify "Numeric" -> True or a column-type spec in ImportString to lock types across chunks.
- For mixed columns that must stay mixed, import as
"RawJSON" or "CSV" with "DataFormat" -> "String" everywhere, then post-process — this avoids silent coercion surprises in either path.
One Missing Diagnostic Detail
What is the typical row count and column count of the window you need to process at once? If your working set fits in < 2 GB of kernel memory, range-constrained Import on a staged file is simpler and more robust. If you must process > 50 M rows incrementally without ever holding more than ~100k rows, the streaming path with explicit type specs is the only viable option.