systemd journal and Linux kernel PSI integration: configuring useful alerts without notification noise
0 reputation · 27 Nov 2024, 01:19 UTC
The goal is to set up alerts that fire when sustained memory, CPU or I/O pressure appears, using the pressure stall information (PSI) files exposed by the kernel and forwarded to the systemd journal, while avoiding alert noise from short‑lived spikes.
PSI reports some, full and avg10 stall‑time fractions; choosing which metric to watch and what avg10 threshold to apply determines sensitivity. The aggregation window used by the alerting pipeline (e.g., node_exporter textfile or cgroup collector) influences how bursty workloads affect the signal, and journald rate limiting (RateLimitIntervalSec, RateLimitBurst) can silently drop rapid pressure events before they reach downstream consumers.
Uncertainty remains about whether to base alerts on PSI some, full, or both, and how to tune the avg10 threshold and aggregation interval for a given workload without generating false positives or missing genuine pressure.
Should alerts be based on PSI some, full, or both, and what avg10 threshold balances sensitivity and noise?
How can the aggregation window be selected to avoid burst‑induced false positives while still detecting sustained pressure?
What journald rate‑limit settings ensure pressure events are not silently dropped before reaching alert pipelines?