Direct recommendation
Use a rate over a short smoothing window, compare to a recent baseline with relative deviation, and require a sustained breach. Do not alert on raw stub_status counters.
Thresholding strategy
Signal-to-noise is best with a combination:
- Rate, not absolute counter. stub_status exposes cumulative counters accepts, handled, requests. Alerting on the absolute value will always fire after the first scrape.
- Relative deviation from a recent baseline for spikes and drops, with an absolute floor for a complete stall.
- Moving average / percentile smoothing instead of instant point values.
Typical PromQL shape:
rate(nginx_requests_total[2m])
Compare to baseline of the same hour-of-day / day-of-week over the last 7-14 days, e.g. p50 or p90 of the same window. Alert when current rate deviates by e.g. -70% for drops or +200% for spikes for a sustained period.
Window size with 1s scrape
With a 1s scrape interval:
- Rate window: 1m to 2m. 1m smooths sub-second bursts while keeping ~60 samples. 2m is safer for bursty clients and scheduled jobs.
- Evaluation for: 2m to 5m. Require the condition to hold for multiple evaluations to suppress transient churn.
Shorter windows catch rapid degradations faster but increase noise. Longer windows suppress noise but delay detection.
Absolute vs relative
Use both:
- Relative deviation for normal variance: catches anomalies in high and low traffic periods without retuning thresholds.
- Absolute floor for critical stalls: rate == 0 or rate < X rps for several minutes indicates nginx is not serving, regardless of baseline.
Confirmed facts vs likely causes
Confirmed
- stub_status metrics are cumulative counters, not gauges.
- Comparing the raw counter to a fixed threshold triggers constantly during normal operation.
Likely explanation for your false positives
- Alert rule uses raw counter or per-second rate without smoothing.
- Threshold set without baseline traffic analysis, so normal diurnal peaks and scheduled jobs fire alerts.
- Scrape interval is short relative to metric granularity, making transient spikes look like sustained anomalies.
Steps needed for this case
- Inspect the alert rule expression to confirm it uses rate() or increase() over a window, not the raw counter.
- Check Prometheus scrape config for the stub_status job interval and compare it to alert evaluation interval and for() duration.
- Plot rate(nginx_requests_total[2m]) over 7 days to establish baseline p50/p95 per hour-of-day and verify the threshold sits above normal burst behavior.
- Replace static thresholds with relative deviation from that baseline and add for: 2m-5m.
One missing diagnostic detail that changes the recommendation: what is the current PromQL alert expression and the scrape interval for the stub_status job? The exact rate window and for duration depend on those values.