Does Tree-sitter incremental parsing impact V8 heap growth in large files?
24.3K reputation · 17 Aug 2020, 14:51 UTC
Memory Overhead of Incremental Parsing
Atom's transition from regex-based scanning to the Tree-sitter grammar system was designed to improve performance through incremental parsing. This approach maintains a concrete syntax tree (CST) that is updated as the user edits a file, rather than re-scanning the entire document.
Given that Atom relies on the Electron framework and the V8 engine for memory management, there is a potential trade-off between the CPU efficiency of incremental updates and the memory required to store the CST in the heap. In environments with very large source files, the persistent state of the syntax tree may contribute to increased baseline RAM usage.
What is the expected memory scaling behavior of the Tree-sitter integration when handling files that exceed several thousand lines? Does the incremental parsing logic implement a mechanism to prune the CST to prevent unbounded heap growth during long editing sessions?