Choosing Prometheus Alert Evaluation Strategies: A Decision Guide
A decision guide for picking Prometheus alert evaluation strategies: immediate firing, duration‑based "for" clauses, and recording rules. Includes a comparison table, trade‑off analysis, and a step‑by‑step implementation you can validate today.
12 Jul 2025, 11:53 UTC

Decision: Choose an Alert Evaluation Strategy
Prometheus can fire alerts immediately when a rule expression becomes true, or it can require the condition to persist for a configurable duration using the for clause. The choice directly affects alert noise, detection latency, and server load. This guide walks through the constraints, compares the supported options, explains trade‑offs, and shows a concrete implementation you can validate today.
Constraints to Consider
- Noise tolerance: How many transient spikes can your on‑call team absorb?
- Detection latency: Is a few‑minute delay acceptable for the monitored symptom?
- Rule evaluation cost: Expensive expressions evaluated every 15 s can degrade Prometheus performance.
- Alertmanager integration: Grouping, inhibition, and silencing depend on the alerts that actually reach Alertmanager.
Supported Options at a Glance
| Strategy | Syntax | Typical Use Case | Impact on Load |
|---|---|---|---|
| Immediate firing | alert: HighErrorRate
expr: rate(http_requests_total{status=~"5.."}[1m]) > 0.05 | Critical, low‑volume signals where any occurrence warrants paging | Low per‑evaluation cost, but can generate many short‑lived alerts if the metric fluctuates |
Duration‑based (for) | alert: HighErrorRate
expr: rate(http_requests_total{status=~"5.."}[1m]) > 0.05
for: 5m | Noisy metrics, capacity‑related thresholds, or any condition that should be sustained | Same evaluation frequency; the for state is kept in memory, negligible extra CPU |
| Recording rule + immediate | groups:
- name: recording
rules:
- record: job:http_error_rate:rate5m
expr: rate(http_requests_total{status=~"5.."}[5m])
- name: alerting
rules:
- alert: HighErrorRate
expr: job:http_error_rate:rate5m > 0.05
for: 2m | Complex or high‑cardinality expressions evaluated once per scrape interval | Reduces per‑evaluation work dramatically; recording rule runs at scrape interval (e.g., 30 s) instead of every rule evaluation (default 15 s) |
Trade‑offs
Immediate vs. for
Immediate alerts fire the moment the expression matches. This is ideal for "hard" failures (e.g., a service health endpoint returns 500). The downside is alert flapping: a brief network blip can create a page that resolves seconds later, increasing on‑call fatigue.
Adding for: 5m tells Prometheus to keep the alert in a pending state until the condition holds continuously for five minutes. The alert only becomes firing after that window, so transient spikes are suppressed. The cost is a detection delay equal to the for duration.
Recording Rules
When an alerting expression contains heavy aggregation (e.g., sum by (job,le) (rate(http_request_duration_seconds_bucket[5m]))), evaluating it every 15 s can push prometheus_rule_evaluation_duration_seconds high and increase memory pressure. A recording rule pre‑computes the expression at the scrape interval, turning the alerting rule into a cheap threshold check.
Trade‑off: you add a new time series (the recorded metric) which increases cardinality. Keep label sets small; avoid recording rules that explode label combinations.
Concrete Implementation: Add a for Clause and a Recording Rule
Assume you monitor HTTP 5xx error rates per job and want to alert only when the rate stays above 5 % for two minutes. The current rule fires immediately.
1. Edit the rule file
# /etc/prometheus/rules/http_errors.yml
groups:
- name: http-error-recording
interval: 30s # matches scrape interval
rules:
- record: job:http_error_rate:rate5m
expr: |
sum by (job) (
rate(http_requests_total{status=~"5.."}[5m])
) /
sum by (job) (
rate(http_requests_total[5m])
)
- name: http-error-alerting
rules:
- alert: HighErrorRate
expr: job:http_error_rate:rate5m > 0.05
for: 2m
labels:
severity: critical
annotations:
summary: "High 5xx error rate for {{ $labels.job }}"
description: "{{ $labels.job }} has a 5xx error rate above 5% for 2 minutes."
Where to run: On the Prometheus server (or the config management host that writes /etc/prometheus/rules/). Requires write permission to the rule directory and the ability to reload Prometheus.
2. Validate syntax
promtool check rules /etc/prometheus/rules/http_errors.yml
Expected output: Checking /etc/prometheus/rules/http_errors.yml
SUCCESS: 2 rules found. If errors appear, fix them before proceeding.
3. Reload Prometheus configuration
# Option A: SIGHUP (requires root or same user as Prometheus process)
kill -HUP $(cat /var/run/prometheus.pid)
# Option B: HTTP POST (if --web.enable-lifecycle is set)
curl -X POST http://localhost:9090/-/reload
Permissions: The user must signal the Prometheus process or reach the HTTP endpoint (typically localhost). Mis‑configured --web.enable-lifecycle exposes a reload endpoint; restrict with firewall or authentication.
4. Verify the recording rule appears
curl -s "http://localhost:9090/api/v1/query?query=job:http_error_rate:rate5m" | jq .data.result
You should see a series per job label with a current value.
5. Confirm alert behavior
# Wait at least 2 minutes after the condition becomes true, then query Alertmanager
amtool alert query alertname=HighErrorRate
If the condition persisted, the alert shows state: firing. If it was transient, the alert stays in pending or does not appear.
Validation Checklist
- Rule evaluation latency: Watch
prometheus_rule_evaluation_duration_seconds{group="http-error-alerting"}in Grafana; it should stay well below the evaluation interval (15 s default). - Recording rule freshness: Query
prometheus_rule_evaluation_duration_seconds{group="http-error-recording"}and ensure the recorded metric updates every scrape interval. - Alertmanager sync: Check
prometheus_alertmanager_discovery_sync_time_secondsto confirm Alertmanager receives the new alert definition. - Cardinality guard: Run
curl -s http://localhost:9090/api/v1/label/__name__/values | jq '.data[] | select(. | startswith("job:http_error_rate"))' | wc -l– the count should equal the number of distinctjoblabels, not explode.
Limitations and Practical Checks
- The
forclause only suppresses firing; the alert remains in pending state and still consumes a small amount of memory per active series. - Recording rules add a new metric name; if you later change the expression, you must also update the recording rule name to avoid stale series.
- Time‑zone handling: Alert timestamps are UTC. If your on‑call schedule uses local time, configure Alertmanager’s
timezonein route/receiver or convert in notification templates. - High‑frequency rule groups (interval < 15 s) increase CPU; keep the default unless you have a measured need.
To verify the whole pipeline end‑to‑end, inject a synthetic 5xx spike (e.g., curl -s -o /dev/null -w "%{http_code}" http://localhost:8080/fail in a loop for three minutes) and confirm the alert transitions from pending to firing exactly after the for window.
Diagram Labels
- Rule Evaluation
- Recording Rule
- Alertmanager Routing
- Metric Scrape
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.