What memory overhead does PromQL query evaluation impose on Prometheus instances?
0 reputation · 11 Aug 2023, 17:46 UTC
Query Memory Overhead Calculation
Prometheus query caching (queries.cache.enabled) is documented as a mechanism to reduce memory pressure during repeated query execution, yet the exact memory overhead calculation for PromQL query evaluation remains undocumented. This creates challenges for capacity planning when deploying complex queries that process high-cardinality metrics.
The tsdb storage engine's memory behavior has evolved significantly across versions, with WAL compaction and retention policies (storage.retention.time) being primary controls for preventing memory growth. However, the relationship between query complexity, result set size, and actual Go runtime heap consumption is not precisely specified in official documentation.
Capacity planning for Prometheus servers requires understanding how different query patterns impact memory usage, particularly when evaluating expressions that involve joins, aggregations, or subqueries across multiple time series.
Which specific PromQL constructs contribute most significantly to memory overhead during evaluation, and how can administrators estimate the required memory capacity for a given query workload?
Can the documented query cache mechanism provide reliable bounds for memory consumption prediction, or does it introduce additional uncertainty in capacity planning scenarios?
How does the memory behavior differ between cached and non-cached query execution paths in terms of Go heap allocation patterns?