Processing Large CSVs with pandas.read_csv Chunksize: Memory‑Efficient Ingestion and On‑the‑Fly Aggregation
Learn how to use pandas.read_csv(..., chunksize=N) to process large CSV files incrementally, keep RAM low, aggregate on the fly, and avoid common pitfalls.
22 Jul 2026, 08:50 UTC

Why chunksize Matters
When a CSV file grows beyond a few hundred megabytes, loading it into a single pandas DataFrame can exhaust RAM and slow the entire pipeline. The chunksize parameter in pd.read_csv turns the reader into a generator that yields small, independent DataFrames. This pattern keeps memory usage bounded while still giving you the full flexibility of pandas for cleaning, transforming, and aggregating data.
How It Works – A Minimal Example
Below is a step‑by‑step illustration of reading a 10‑million‑row CSV, computing a column sum on the fly, and verifying the result. Replace YOUR_CSV_PATH with the actual file location.
import pandas as pd
import os
csv_path = "YOUR_CSV_PATH"
# Choose a chunk size that comfortably fits in memory – 100,000 rows is a common starting point.
chunksize = 100_000
# Accumulator for the running total of the numeric column.
column_sum = 0
# Open the CSV as an iterator of DataFrames.
for i, chunk in enumerate(pd.read_csv(csv_path, chunksize=chunksize, usecols=["value"], dtype={"value": "int32"} )):
# Each chunk is a DataFrame with at most rows.
# Perform any per‑chunk processing here; we only sum for illustration.
chunk_sum = chunk["value"].sum()
column_sum += chunk_sum
# Optional: print progress to track long runs.
if (i + 1) % 10 == 0:
print(f"Processed { (i + 1) * chunksize } rows – running sum: {column_sum}")
print(f"Final sum of 'value' column: {column_sum}")
Key points in the code:
chunksize=100_000creates an iterator that yields DataFrames of up to 100,000 rows.- Using
usecolsanddtypereduces parsing time and memory footprint. - The loop keeps a
column_sumvariable that accumulates results across chunks. - Because the loop discards each chunk after processing, the peak memory stays near the size of a single chunk plus overhead.
Practical Verification
After running the script, you can confirm correctness by comparing to a full‑file read (only for sanity checks, not in production). For example:
full_df = pd.read_csv(csv_path, usecols=["value"], dtype={"value": "int32"})
assert full_df["value"].sum() == column_sum, "Mismatch detected"
print("Verification passed – sums match")
To monitor memory usage, tools like memory_profiler or tracemalloc can be attached around the loop. A typical peak will be roughly the memory of one chunk plus a small overhead, regardless of the total file size.
Choosing the Right Chunksize
There is no one‑size‑fits‑all value. The optimal chunksize depends on:
- Available RAM and the size of a single chunk in memory.
- I/O bandwidth – very small chunks mean more file reads and higher overhead.
- Processing complexity – heavy per‑chunk transformations can justify larger chunks to amortize loop costs.
A practical approach is to benchmark a few sizes (e.g., 10k, 100k, 1M rows) and plot runtime vs. memory usage. The sweet spot is where runtime is acceptable and peak memory stays well below the system limit.
Common Mistakes and Edge Cases
- Assuming order is preserved. If you shuffle or sort within a chunk, the overall order may become scrambled unless you explicitly merge and re‑sort after the loop.
- Using groupby or other full‑DataFrame operations after the loop. These require the entire dataset in memory. If you need a global aggregation, perform it incrementally (e.g., maintain a dictionary of group totals) or concatenate the chunks at the very end if the final size is manageable.
- Skipping rows with
skiprowsandchunksize. The first chunk may inadvertently include header lines ifskiprowsis not configured correctly. Test that the header is read only once. - Mixed dtypes across chunks. If a column contains values of different types in different chunks, pandas may upcast the dtype, increasing memory usage. Explicitly set
dtypefor columns to avoid this. - Writing without appending. If you write processed chunks to a new file, use
mode='a'andheader=Falseafter the first write to preserve all data.
When to Concatenate All Chunks
Sometimes you need a single DataFrame for downstream libraries that only accept a DataFrame input. In that case, accumulate the processed chunks in a list and call pd.concat once:
processed_chunks = []
for chunk in pd.read_csv(csv_path, chunksize=chunksize):
processed_chunks.append(chunk_processed)
full_df = pd.concat(processed_chunks, ignore_index=True)
Be mindful that this defeats the memory advantage if the final DataFrame is still large. Only use this pattern when the final size is within available RAM.
Limitations
- The iterator does not support random access – you cannot seek to a specific row without re‑reading from the start.
- Some options, like
na_filterorconverters, are applied per chunk and may produce slightly different results than a full read if the data contains edge cases. - Large
chunksizevalues may cause a single chunk to exceed memory, especially when many columns are read with default dtypes.
Bottom Line
Using pd.read_csv(..., chunksize=N) is the most straightforward way to process massive CSV files with pandas while keeping memory usage predictable. By choosing a sensible chunk size, specifying usecols and dtype, and performing per‑chunk aggregation, you can ingest and analyze datasets that would otherwise be out of reach on commodity hardware.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.