Query vs Eval: Preventing Stale pandas Results
25.5K reputation · 11 Nov 2022, 03:52 UTC
Background
When repeatedly executing DataFrame.query on the same DataFrame, pandas compiles the query string into a numexpr expression and caches the compiled Abstract Syntax Tree (AST). The cache is keyed solely on the query string, not on the underlying data. Consequently, if the DataFrame is mutated after the first evaluation, subsequent query calls may return stale results that reflect the original state.
Two Documented Work‑arounds
1. DataFrame.query(..., engine='python') disables the numexpr engine entirely, forcing a fresh Python evaluation on each call. This eliminates stale reads but incurs higher CPU usage and slower performance on large datasets.
2. DataFrame.eval(..., local_dict=…) forces a new evaluation by supplying a local dictionary, thereby bypassing the numexpr cache. It retains the numexpr engine for speed but requires the user to provide the dictionary, which can increase memory consumption if the dictionary is large.
Unresolved Decision
Both approaches avoid stale cache but trade off speed, memory usage, and code complexity. The choice depends on the frequency of DataFrame updates, dataset size, and resource constraints.
- How does the numexpr cache invalidate (or not) when a DataFrame is modified?
- What is the memory overhead of using
evalwith a local dictionary compared to disabling numexpr inquery? - In a production environment with frequent DataFrame updates, which approach offers the best balance of correctness, performance, and memory usage?