Question
Pandas read_csv low_memory flag causes unexpected dtype promotion and memory spikes
Tasadduq BurneyownerOwner · Founder
25.5K reputation · 14 Mar 2021, 18:29 UTC
13.8K views0
Observed behavior
When loading a CSV with low_memory=False, columns that contain only integers but a few NaNs are promoted to float. The resulting DataFrame consumes noticeably more memory than expected, even though the flag is intended to reduce peak usage.
Unresolved questions
- Does
low_memory=Falseguarantee a single‑pass read, or does pandas still perform a second pass for dtype inference? - How does the dtype inference heuristic treat mixed‑type columns when
low_memory=Falseis set, especially regarding integer‑to‑float promotion? - Is there a documented version change between pandas 1.3 and 2.x that alters the memory allocation strategy for
low_memory=False?