Cut Dashboard Latency with Prometheus Recording Rules: A Practical Guide
Cut Grafana panel load times by pre‑computing heavy PromQL expressions with Prometheus recording rules. Learn how to set them up, verify them, and balance storage and CPU costs.
11 Feb 2026, 02:58 UTC

Why Your Dashboards Are Slow
When a Grafana panel takes seconds to render, the culprit is almost always a PromQL expression that forces Prometheus to scan millions of raw samples and perform heavy aggregations. Common patterns that trigger this include sum_over_time(), rate() over long windows, or multi‑label reductions on metrics that have dozens of unique label combinations. The result is a high CPU load on the Prometheus server and a poor user experience.
Recording Rules: Move the Work to Ingestion Time
A recording rule tells Prometheus to pre‑compute a PromQL expression at a regular interval and store the result as a new time series. Subsequent queries can then fetch that pre‑computed series instead of re‑calculating the expression each time.
Key points:
- Evaluation happens at the
evaluation_intervaldefined in the global section ofprometheus.ymlor overridden per rule group. - Each evaluation appends a new sample to the recorded series, so the rule must be written with the desired resolution in mind.
- Recording rules do not back‑fill existing data; they start collecting from the moment the rule is loaded.
- They can be defined in the same file as alerting rules; a reload will re‑evaluate all rules in the file.
Setting Up a Recording Rule
Below is a minimal example that pre‑computes the 5‑minute request rate for an API service. The rule will be evaluated every 30 seconds.
# /etc/prometheus/rules/api-recording.yml
groups:
- name: api-recording
interval: 30s
rules:
- record: job:http_requests:rate5m
expr: sum(rate(http_requests_total{job="api-server"}[5m]))
Steps to deploy:
- Place
api-recording.ymlon a directory that the Prometheus process can read. - Reference the file in
prometheus.yml:rule_files: - "rules/api-recording.yml" - Reload the configuration. If the
/-/reloadendpoint is enabled and authenticated, you can use:
Otherwise send acurl -X POST http://localhost:9090/-/reloadSIGHUPto the Prometheus PID.
Verifying the Rule Works
Prometheus exposes the status of all rules at /api/v1/rules or the Rules tab in the web UI. A healthy rule will show OK and a recent last_evaluation timestamp.
To compare performance, run the original expression and the recorded metric in the Expression Browser:
sum(rate(http_requests_total{job="api-server"}[5m]))job:http_requests:rate5m
Both queries should return the same value, but the recorded metric will typically respond faster because it is a simple lookup. Keep in mind that the freshness of the recorded value depends on the evaluation interval; a 30s interval means the data can be up to 30 seconds out of date.
Trade‑offs and Practical Limits
While recording rules reduce query latency, they come with trade‑offs that you should monitor:
- Storage growth: Each rule creates an additional series. If the rule retains high‑cardinality labels (e.g.,
pod), the TSDB will store a separate series for every label combination, quickly inflating disk usage. - CPU overhead: The rule evaluation itself consumes CPU proportional to the complexity of the expression and the number of series it touches. A very frequent evaluation interval can become a bottleneck.
- Data freshness: The recorded series lags behind raw data by up to one evaluation interval. For metrics that need millisecond precision, recording rules may not be suitable.
- Rule conflicts: If a recording rule and an alert rule share the same
recordname, the alert rule will take precedence and the recorded series will not be created. - No retroactive data: Rules start collecting from the moment they are loaded; they do not compute values for historical data.
Mitigation strategies include:
- Drop unnecessary labels from the expression using
label_replaceorwithoutclauses. - Use a moderate evaluation interval (e.g., 30s-1m) that balances CPU usage and data freshness.
- Monitor
prometheus_tsdb_storage_size_bytesandprometheus_tsdb_head_seriesmetrics to detect runaway growth. - Keep rule names unique and follow a convention like
job:metric:operationto avoid accidental overrides.
Actionable Checklist
- Identify the slowest PromQL queries by checking
/api/v1/queryresponse times or using Grafana’s query inspector. - For each candidate query, evaluate whether the result can be pre‑computed and whether the series will maintain acceptable cardinality.
- Create a recording rule with an appropriate evaluation interval.
- Reload Prometheus and verify the rule status.
- Replace the original query in dashboards with the recorded metric and observe the latency improvement.
- Continuously monitor TSDB size and CPU load; adjust the evaluation interval or drop labels if needed.
By following this workflow, you can systematically reduce dashboard latency while keeping an eye on storage and CPU costs.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.