Diagnosing Grafana Alert Rules That Fail to Fire
Learn how to identify why a Grafana alert rule does not trigger even when its condition is met, with a step‑by‑step checklist, common causes, and fixes.
04 Aug 2026, 17:10 UTC

Recognizable condition
You have configured an alert rule in Grafana, but the rule never enters the firing state even though the underlying metric clearly satisfies the condition (e.g., CPU usage > 80 %). No notifications are sent, and the rule appears inactive in the alert list.
Common causes and quick diagnostic table
| Symptom | Likely cause | Where to look |
|---|---|---|
| Rule does not appear in the Alerting > Rules list | Rule failed to load (YAML syntax error, missing query/condition) | Grafana UI > Alerting > Rules; check for error banners |
| Rule shows “No data” or “Error” state | Underlying query returns no results or times out | Alert rule > Evaluation tab; inspect query output |
| Rule stays in “Pending” state longer than expected | Evaluation interval too long or ‘for’ duration exceeds interval | Alert rule > Condition tab; review interval and ‘for’ fields |
| Rule evaluates correctly but no notification arrives | Notification channel misconfigured, disabled, or silenced | Alerting > Notification channels; test channel; Alerting > Silences |
| Rule fires but history is missing after restart | Alert state storage backend lacks retention policy | Grafana configuration > alerting.storage; check backend (e.g., InfluxDB) retention |
Ordered checks
- Verify rule loading
- Where: Grafana UI > Alerting > Rules
- Permissions: Any user with Alerting Read access; admin needed to edit
- Check: Look for a red error banner or “Failed to load” message.
- Risk: None – read‑only operation.
- Inspect rule definition
- Where: Open the rule > click “Edit”
- Permissions: Alerting Edit
- Check: Ensure the
queryfield references a valid data source and returns a numeric series; verify the condition operator (above, below, missing) matches intent. - Risk: None if you only view; editing incorrectly could break the rule.
- Preview query results
- Where: Alert rule > Evaluation tab
- Permissions: Alerting Edit
- Check: Select a recent time range (e.g., last 15 minutes) and press “Run”. Confirm that the query returns values and that the condition evaluates to true for at least one sample.
- Risk: None – read‑only.
- Check evaluation interval and ‘for’ duration
- Where: Alert rule > Condition tab
- Permissions: Alerting Edit
- Check: Ensure the evaluation interval is ≤ the scraping interval of the data source (e.g., if Prometheus scrapes every 30 s, interval should be 30 s or multiples). Verify that the
forduration is ≥ evaluation interval; aforshorter than the interval is ignored and may cause unexpected behavior. - Risk: Changing interval may increase load on the data source; test in a staging environment first.
- Test notification channel
- Where: Alerting > Notification channels > select channel > “Test”
- Permissions: Alerting Admin (to edit channels) or Alerting Read (to run test if already configured)
- Check: You should receive a test message (Slack, email, etc.). If the test fails, note the error (e.g., invalid webhook URL, authentication failure).
- Risk: None – test only sends a message.
- Look for active silencers
- Where: Alerting > Silences
- Permissions: Alerting Read
- Check: Any silence that matches the rule’s labels will suppress notifications. Note the silence expiry.
- Risk: None.
- Verify alert state storage (if using external backend)
- Where: Grafana configuration file (
grafana.ini) >[alerting]section or via UI > Settings > Alerting > Storage - Permissions: Grafana Admin
- Check: Confirm that
enabled = trueand a valid backend (e.g., InfluxDB, OpenTSDB) is configured. Then query the backend for recent alert state entries (e.g.,SELECT * FROM alert_state WHERE time > now() - 1h). - Risk: Querying the backend may affect performance; use a limited time range.
- Where: Grafana configuration file (
Fixes tied to findings
- Rule failed to load – Correct YAML syntax: ensure required keys
title,condition,queryare present and properly indented. Reload the rule after saving. - Query returns no data – Validate the data source query independently (Explore view). Adjust the query time range or add missing labels.
- Evaluation interval mis‑set – Set interval to a multiple of the data source scrape interval (e.g., 30s, 1m). Avoid intervals shorter than the scrape interval.
- ‘for’ duration too short/long – Set
forto at least one evaluation interval; adjust based on desired alerting latency (e.g., 5 minutes for a 1 minute interval). - Notification channel broken – Fix the channel configuration (correct webhook URL, credentials, enable the channel). Re‑run the test.
- Active silencer – Either delete the silence or wait for it to expire; ensure silences are applied only when intentional maintenance occurs.
- Missing alert state retention – Configure a retention policy on the backend (e.g., InfluxDB
RETENTION POLICY \"alerts\" ON \"grafana\" DURATION 30d REPLICATION 1) and restart Grafana.
Escalation criteria
- After completing all checks, the rule still shows “No data” or “Error” despite a valid query in Explore.
- Notification channel tests succeed but alerts never transition to firing, even after reducing
forto zero. - Alert state storage backend shows no entries being written, indicating a deeper storage or plugin issue.
In these cases, enable Grafana debug logging ([log] level = debug) and reproduce the condition. Capture the logs around the alert evaluation timestamp and share them with Grafana support or the community, including Grafana version, data source type, and the exact alert rule YAML.
Example: CPU usage alert rule
# alert-rule.yaml
apiVersion: 1
groups:
- name: system-metrics
interval: 1m
rules:
- alert: HighCPUUsage
expr: |
avg by (instance) (rate(node_cpu_seconds_total{mode=\"system\"}[5m])) > 0.8
for: 5m
labels:
severity: critical
annotations:
summary: \"CPU usage above 80% on {{ $labels.instance }}\"
description: \"CPU has been above 80% for more than 5 minutes.\"
To verify this rule:
- Import the YAML via Alerting > Rule management > Import.
- Open the rule, go to the Evaluation tab, select the last 10 minutes, and run the query. You should see a time series with values > 0.8 when CPU is high.
- Check that the interval is 1 m (matches Prometheus scrape) and the
foris 5 m (≥ interval). - Test the notification channel (e.g., Slack) and confirm you receive a test message.
- After waiting at least 5 minutes of sustained high CPU, the rule should move to firing state and trigger the notification.
Limitations and practical checks
- The evaluation interval cannot be shorter than the underlying data source’s scrape interval; otherwise the query will return stale or no data.
- The
forfield is ignored if it is less than the evaluation interval; Grafana treats it as zero, which may cause immediate firing. - If using the default built‑in alerting storage (SQLite or MySQL), alert state is persisted locally; ensure the
datadirectory is backed up and has sufficient disk space. - Practical check: after any change, use the Evaluation tab to confirm the rule’s condition evaluates to true, then wait for at least one evaluation interval plus the
forduration before expecting a notification.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.