Answer to the Core Question
When you repeatedly open a very large pandas.DataFrame in DataSpell, the IDE’s built‑in preview keeps a reference to the original object in its internal cache until the preview tab is closed. Each new preview adds another reference, so the memory footprint grows linearly with the number of opened tabs. In contrast, pandas‑profiling (now pandas-profiling) generates a static HTML report by consuming the entire DataFrame in a single run, then discards the DataFrame reference once the report is written. Therefore, for repeated opening of the same large DataFrame, DataSpell’s preview typically uses less memory if you close the tab after each use; if you keep tabs open, the cumulative memory can exceed what the profiling tool would use for a single run.
Programmatic Cache Clearing in DataSpell?
DataSpell provides no documented API or runtime command to evict the preview cache for a specific tab. The only built‑in mechanism is the global Invalidate Caches / Restart command, which clears the entire IDE cache but is disruptive for interactive work. Consequently, the recommended practice is to close the preview tab (or use a small preview size) when you no longer need it.
Disk‑Usage & Reproducibility Trade‑offs of pandas‑profiling
- Disk Usage: The generated report is an HTML file that can be several megabytes for very large DataFrames. It is a one‑time write; subsequent openings of the report do not re‑read the DataFrame, so disk I/O is minimal after the first run.
- Reproducibility: The report is static. If the underlying DataFrame changes, the report does not update automatically, so you must re‑run
ProfileReport to capture the new state. This guarantees a snapshot that can be shared or archived.
- Memory During Generation: The profiling library holds the full DataFrame in memory while computing statistics, which can lead to
MemoryError for very large tables unless you use the sample parameter or pre‑filter the data.
Practical Steps to Compare Memory Footprints
Load the large CSV into a pandas.DataFrame in a fresh Python session.
Measure peak memory with tracemalloc or memory_profiler while:
- Opening the DataFrame in DataSpell’s preview (ensure the preview size is set to its default, e.g., 200 rows).
- Running
ProfileReport(df).to_file('report.html') in a separate script.
Compare the peak memory values. For a DataFrame of 10 M rows, you’ll typically see ~200 MB for the preview (depending on preview size) versus ~1 GB for the profiling run.
Bottom Line
• If you can close preview tabs promptly, DataSpell’s preview is lighter on memory and avoids the full‑DataFrame load of pandas‑profiling. • If you need a persistent, shareable snapshot or cannot keep tabs open, generate a pandas‑profiling report; just be mindful of the temporary memory spike and consider sampling for very large datasets. • No programmatic cache eviction exists in DataSpell; rely on tab closure or IDE cache invalidation.