Graceful Shutdown Under Memory Pressure with cgroup v2 PSI Notifications
Use cgroup v2 PSI memory.pressure triggers or systemd MemoryPressureWatch to start a graceful shutdown while your service still has memory to clean up — before the OOM killer decides for you.
06 Oct 2026, 23:44 UTC

The OOM killer doesn't send a warning letter
When a Linux service runs out of memory, the usual ending is abrupt: the kernel's out-of-memory (OOM) killer picks a victim and sends SIGKILL. No cleanup, no connection draining, no checkpoint. For a database, queue worker, or anything holding in-flight state, that's the worst possible way to stop.
The useful takeaway: cgroup v2 exposes memory pressure information through the Pressure Stall Information (PSI) interface, and applications (or systemd) can watch it. Instead of waiting for the hard limit to trigger a kill, your process can observe sustained memory reclaim difficulty and start a graceful shutdown while it still has working memory to do so.
What PSI actually tells you
PSI, available since kernel 4.20, doesn't report "bytes used." It reports stall time: how long tasks were delayed waiting for memory. Each cgroup v2 directory has a memory.pressure file with two lines:
some avg10=0.00 avg60=0.00 avg300=0.00 total=0
full avg10=0.00 avg60=0.00 avg300=0.00 total=0some means at least one task stalled on memory; full means all non-idle tasks stalled simultaneously. avg10 is the percentage of wall-clock time stalled over the last 10 seconds. High some values mean the kernel is actively reclaiming (swapping, dropping page cache) to satisfy allocations — the early-warning phase before an OOM kill.
This is a better shutdown trigger than a raw usage threshold like "90% of memory.max," because usage alone says nothing about distress. A process can sit at 95% of its limit happily if the memory is clean page cache the kernel can drop instantly.
Watching pressure from an application
PSI supports trigger-based notification: you write a threshold to the file, then poll() the descriptor. The kernel wakes you when the stall average crosses your threshold within the given window. A minimal sketch (run inside the cgroup, needs read/write access to its memory.pressure file):
int fd = open("/sys/fs/cgroup/myapp/memory.pressure", O_RDWR | O_CLOEXEC);
/* Trigger when 'some' stalls exceed 50ms in a 1s window */
dprintf(fd, "some 50000 1000000");
struct pollfd pfd = { .fd = fd, .events = POLLPRI | POLLERR };
while (poll(&pfd, 1, -1) > 0) {
if (pfd.revents & (POLLERR | POLLHUP))
break; /* cgroup removed or fd invalidated */
initiate_graceful_shutdown(); /* drain, checkpoint, then exit */
}Thresholds are in microseconds: some 50000 1000000 means "at least 50 ms of partial stall within any 1-second window." Note that registering a trigger requires write permission on the file — in containers, that usually means the cgroup must be delegated to the container (e.g., via a user namespace or systemd's Delegate=yes). A read-only bind of /sys/fs/cgroup lets you sample averages but not register triggers.
Letting systemd do the watching
If you'd rather not write polling code, recent systemd versions (v252 and later; check systemctl --version) can monitor PSI for you:
[Service]
MemoryHigh=1G
MemoryMax=1200M
MemoryPressureWatch=auto
MemoryPressureThresholdSec=200msWith MemoryPressureWatch=auto, systemd monitors the unit's cgroup PSI and, when pressure exceeds the threshold, it acts on the service — by default it can trigger the configured OOM policy for the unit rather than letting the kernel pick a victim system-wide. Verify the settings are applied:
systemctl show myapp.service \
--property=MemoryPressureWatch,MemoryPressureThresholdSecRun this as root or with appropriate polkit permissions. If the properties show defaults despite your unit settings, your systemd version may predate the feature — confirm before relying on it.
A deliberate pairing matters here: set MemoryHigh (the throttling limit, where the kernel reclaims aggressively) below MemoryMax (the hard limit). PSI pressure rises as the workload pushes against MemoryHigh, giving your watcher a window to act before MemoryMax forces an OOM kill.
Trade-offs and failure modes
- False positives. A short, intense reclaim burst — say, a backup job reading a large file into page cache — can trip a tight threshold. Tune against real workload profiles; a 10-second-average-based trigger is less twitchy than a 1-second one.
- Shutdown needs memory too. Graceful draining, checkpointing, and flushing buffers all allocate. If you wait until pressure is critical, the shutdown path itself may get OOM-killed. Trigger early, at moderate sustained pressure.
- Descriptor lifetime. If the cgroup is removed (service stopped, container destroyed), the fd becomes invalid. Handle
POLLERR/POLLHUPand treat it as "stop watching," not a crash. - cgroup v1 systems. PSI triggers are a cgroup v2 feature. Check with
stat -fc %T /sys/fs/cgroup—cgroup2fsmeans v2. On v1 you're limited to polling usage counters, which loses the stall-time signal entirely.
Checking that it works
Don't trust the wiring until you've seen it fire. Put a test service in a cgroup with a low MemoryHigh, then generate pressure — stress-ng --vm 1 --vm-bytes 800M works well. Watch cat /sys/fs/cgroup/<path>/memory.pressure climb, confirm your handler logs a trigger event, and confirm the process exits cleanly with state flushed rather than appearing in memory.events under oom_kill. If oom_kill increments before your handler fires, your threshold is too late — lower it or lower MemoryHigh to widen the window.
The actionable starting point: pick one stateful service, give it a MemoryHigh 15–20% below its hard limit, add a PSI watcher (or MemoryPressureWatch if your systemd is new enough), and rehearse the failure. An OOM kill you engineered in testing is far cheaper than the one production schedules for you.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.