The short answer
There is no graceful transition. When a ZFS pool or a dataset quota runs out of space, writes fail immediately with ENOSPC, regardless of what the guest OS believes is free. Proxmox VE does not intercept this at the storage layer — the guest's virtual disk simply starts returning I/O errors. The only reliable defense is preventing the pool from ever reaching that state: keep real usage well below capacity, monitor utilization, and alert early.
Why the guest sees corruption risk
Thin provisioning means the guest was told it has, say, 500 GB, while the pool may physically hold far less. The guest filesystem has no visibility into pool health, so it happily issues writes that ZFS must reject. Two distinct limits can trigger this:
- Pool exhaustion: ZFS performance and allocation behavior degrade as the pool fills; operationally, treat roughly 80% as the ceiling. Beyond that, copy-on-write allocation struggles to find contiguous free space, and writes can fail even before the pool reports 100%.
- Dataset/volume quota: a quota on a zvol or dataset is enforced at write time. Any write that would exceed it returns
ENOSPC instantly, independent of free space elsewhere in the pool.
Either failure surfaces in the guest as a generic disk error, which filesystems typically handle by remounting read-only or corrupting in-flight data.
What actually protects you
1. Monitor pool usage and alert before the limit. This is the real answer to your second question — the early-warning mechanism is monitoring, not a ZFS setting. Check usage with:
zpool list -o name,size,alloc,free,capacity,fragmentation
zfs list -r -o name,used,avail,quota,reservation,referenced
Wire capacity into your monitoring (Prometheus node exporter textfile, Zabbix, or a cron job that mails when capacity exceeds a threshold). Alert at 70%, act by 80%.
2. Use reservations and quotas deliberately. A refquota on a VM disk volume caps how much that guest can consume, converting an uncontrolled pool-wide failure into a per-VM, predictable one. A reservation on critical datasets guarantees they always have space. Note that snapshots count toward a quota but not toward a refquota — that distinction matters when deciding which to apply.
3. Manage snapshot growth. On an over-provisioned pool, automatic snapshots are the most common cause of sudden exhaustion: deleted-but-snapshotted data still occupies space. Cap snapshot retention and include snapshot usage in your alerts (zfs list -o usedbysnapshots).
4. Size conservatively. Over-provisioning is only safe if the sum of realistic guest growth stays under your alert threshold. If guests genuinely may fill their disks, the pool must physically hold that data.
Verifying behavior before production
You can confirm enforcement semantics on a scratch dataset:
zfs create -o quota=1G tank/testquota
dd if=/dev/zero of=/tank/testquota/fill bs=1M count=2000
The dd will stop with "No space left on device" at the quota boundary — the same error a guest VM would receive. Also verify where a quota comes from with zfs get -r quota tank, since inherited quotas are a common surprise.
One caveat
Exact safe-fill thresholds depend on pool layout (mirror vs. RAIDZ width), record/volblock sizes, and fragmentation, so treat 80% as a conservative rule of thumb rather than a documented hard limit. If you can share your pool layout (zpool status) and whether you use RAIDZ, the threshold and alert levels can be tuned more precisely.