Reducing Dashboard Latency with Prometheus Recording Rules
Stop fighting Prometheus query timeouts. Learn how to use Recording Rules to pre-compute expensive PromQL expressions and drastically speed up your Grafana dashboards.
02 May 2026, 08:17 UTC

The Query Timeout Problem
You build a Grafana dashboard to monitor a high-traffic cluster. Everything looks great when viewing the last 5 minutes of data. However, when you switch the time range to 7 days or 30 days, the panels spin indefinitely before returning a "Query Timeout" or "Memory Limit Exceeded" error.
This happens because Prometheus is a pull-based system that calculates results at query time. If you are summing the rates of thousands of individual pods across a month of data, Prometheus must fetch every single raw sample from the TSDB (Time Series Database) and perform the math on the fly. As your cardinality—the number of unique label combinations—increases, this approach becomes unsustainable.
The solution is to move the computational burden from the read path to the write path using Recording Rules.
What are Recording Rules?
Recording rules allow you to pre-compute a complex PromQL expression and save the result as a new, standalone time series. Instead of calculating a sum or average every time a user refreshes a dashboard, Prometheus calculates it once at a defined interval and stores the result. When you query the recording rule metric, Prometheus treats it like any other raw metric, returning the pre-calculated value instantly.
Implementing a Performance-Optimized Rule
To implement recording rules, you define them in a YAML file and tell the Prometheus server where to find that file via the rule_files configuration block.
Step 1: Define the Rule
Create a file named rules.yml. In this example, we are pre-computing the total request rate across a cluster to avoid calculating a sum(rate(...)) over millions of samples during a dashboard load.
groups:
- name: cluster_performance_rules
interval: 1m
rules:
- record: job:request_rate:sum_total
expr: sum(rate(http_requests_total[5m]))
Step 2: Configure Prometheus
Update your prometheus.yml to include the rules file. This requires a reload of the Prometheus configuration or a restart of the service.
global:
scrape_interval: 15s
rule_files:
- "rules.yml"
Step 3: Verification
To verify the rule is working, run the following checks in the Prometheus expression browser:
- Check for existence: Query
job:request_rate:sum_total. If it returns data, the rule is active. - Compare performance: Run the raw expression
sum(rate(http_requests_total[5m]))and compare the execution time (visible in the Prometheus UI) against the recording rule metric. For large datasets, the difference is often several orders of magnitude.
Naming Conventions and Cardinality
Because recording rules create new metrics, naming is critical to avoid confusion between raw data and pre-computed data. A widely accepted pattern is level:metric:operation.
- level: The scope of the aggregation (e.g.,
job,cluster,instance). - metric: The name of the original metric being processed.
- operation: The math performed (e.g.,
sum,mean,rate).
Following this pattern ensures that anyone looking at the TSDB knows exactly where the data came from and that it is a derived value, not a raw measurement from an exporter.
Trade-offs and Limitations
Recording rules are not a "silver bullet" and come with specific costs:
- Increased Storage: Every recording rule creates a new time series. If you create rules with high-cardinality labels, you are increasing the disk space required for your TSDB.
- No Retroactive Data: Recording rules only start calculating from the moment they are deployed. You cannot use a recording rule to speed up a query for data from last month if the rule was only created today.
- Resolution Lag: The
intervalsetting determines how often the rule is evaluated. If you set a 5-minute interval, your pre-computed metric will only update every 5 minutes, regardless of how often the raw metrics are scraped.
Decision Summary
Use recording rules when a specific PromQL query is used frequently across multiple dashboards or alerts and consistently takes longer than a few seconds to execute. If the query is only run once a week for a report, the storage cost of a recording rule likely outweighs the benefit. For high-traffic production environments, shifting the load to the write path is the most effective way to maintain dashboard responsiveness.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.