Stop Waiting for Grafana: Using Prometheus Recording Rules to Fix Slow Dashboards
Stop battling slow Grafana dashboards. Learn how to use Prometheus recording rules to precompute expensive PromQL expressions and drastically reduce query latency.
17 May 2026, 14:50 UTC

The Dashboard Timeout Problem
You have a high-traffic service with hundreds of pods. Your Grafana dashboard looks great, but it takes 15 seconds to load, or worse, it throws a gateway timeout when you try to view a 7-day trend. The culprit is usually a heavy PromQL aggregation—like a sum by (job) (rate(...))—running across thousands of time series every time the page refreshes.
The solution isn't always adding more CPU to your Prometheus server. Instead, you can use Recording Rules to precompute these expensive calculations and store the results as new, lightweight time series. This shifts the computational burden from the query time (when the user is waiting) to the evaluation time (background processing).
How Recording Rules Work
A recording rule is essentially a PromQL expression that Prometheus evaluates on a fixed schedule. Instead of just triggering an alert, the result is written back into the Time Series Database (TSDB) as a new metric.
When you query this new metric, Prometheus doesn't have to re-calculate the rate or sum across all your pods; it simply reads the pre-aggregated value. This transforms a query that might scan 10,000 samples into one that scans 10.
The Naming Convention
Because recording rules create new metrics, naming is critical to avoid confusion. A widely accepted community standard is level:metric:operations. For example, job:http_requests_total:rate5m tells any engineer looking at the metric that it is aggregated at the job level, based on http_requests_total, and represents a 5-minute rate.
Worked Example: Optimizing Request Rates
Imagine you are monitoring a service where http_requests_total is scraped from 500 different pods. A standard dashboard query looks like this:
sum by (job) (rate(http_requests_total[5m]))
To optimize this, create a rule file (e.g., rules.yml) with the following configuration:
groups:
- name: http_request_aggregations
interval: 1m
rules:
- record: job:http_requests_total:rate5m
expr: sum by (job) (rate(http_requests_total[5m]))
Implementation Details:
- Run Location: This file is loaded by the Prometheus server via the
--rule.filesflag. - Permissions: The Prometheus process must have read access to the YAML file.
- Expected Result: A new metric named
job:http_requests_total:rate5mwill appear in your expression browser.
Now, update your Grafana panel to query job:http_requests_total:rate5m directly. The dashboard will load nearly instantaneously because the heavy lifting is happening once per minute in the background, rather than every time a user hits refresh.
The Engineering Trade-offs
Recording rules are not a "magic button"; they introduce specific costs and limitations:
- Storage Overhead: You are creating new time series. While usually smaller than the raw data, high-cardinality recording rules can increase TSDB disk usage and memory pressure.
- Resolution Loss: The recorded series only has data points at the
intervaldefined in the rule group (e.g., every 1 minute). You cannot use a recording rule for forensic analysis that requires the original scrape granularity (e.g., 15-second intervals). - Data Persistence: If you deploy a rule with a logic error, the incorrect data is written to the database. Fixing the rule corrects future data, but the historical samples remain wrong.
Verification and Safety
Before deploying rules to production, use the promtool CLI (shipped with Prometheus) to validate your syntax. Run this command on your local machine or in a CI pipeline:
promtool check rules rules.yml
Once deployed, verify the health of your rules by checking these internal Prometheus metrics:
| Metric | What to look for |
|---|---|
prometheus_rule_group_last_duration_seconds |
Ensure the evaluation time is significantly lower than the rule interval. |
prometheus_rule_group_iterations_missed_count |
Any value above 0 indicates the server is overloaded and skipping evaluations. |
Rollback Procedure
To remove a recording rule, remove the rule from the YAML file and restart or reload the Prometheus configuration. Note that this stops the generation of new data, but the existing recorded samples will remain in the TSDB until they hit the global retention limit.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.