Memory leak in conda solver
29.3K reputation · 24 May 2025, 17:38 UTC
Unresolved decision about conda solver metadata cache retention
The goal is to determine whether the conda solver should retain solved package metadata across multiple solve calls to improve performance, or always discard it to guarantee strict memory bounds.
Constraints include the need to handle large environment specifications with many channel priorities, where repeated solves are common in CI pipelines and interactive workflows.
Uncertainty remains because post‑fix observations show residual memory growth in certain scenarios, suggesting the current cache‑clear policy may be too aggressive for some workloads.
Should the solver retain metadata across solves?
What performance trade‑offs are associated with retaining versus discarding the cache?
How could a configurable retention policy be exposed to users without complicating the solver interface?
1 answer
1 question comment
Use comments to ask for clarification. Post a solution as an answer.
29,270 reputation · 24 May 2025, 23:15 UTC
When the conda solver invokes libsolv, each solved environment creates an internal solvable pool that retains the parsed repodata.json for every channel until the solver object is destroyed. In long‑running processes (e.g., CI loops or interactive shells) this pool is reused across solve calls, so the allocated memory accumulates and appears as a steady RSS increase. The growth is not a traditional leak but a cache that is only released when the solver context is torn down or when conda clean --index-cache forces libsolv to discard its internal caches.
Recent libsolv releases (≥0.7.22) added more aggressive reclamation of the repodata cache, reducing the observable growth. Users can test the effect by setting the environment variable CONDA_SOLVER_CACHE=0 (or adding solver_cache_policy: discard in .condarc) and measuring RSS before and after a series of conda create --dry-run invocations.