Using Proxmox VE’s Native ZFS Replication Scheduler for Warm Standby DR
Configure Proxmox VE’s native ZFS replication scheduler to create point‑in‑time VM and container copies for low‑RPO warm standby and off‑site disaster recovery.
20 Aug 2026, 03:14 UTC

Problem: Need a low‑RPO warm standby for VMs and containers
In a small‑to‑medium Proxmox VE cluster you often want a copy of critical VMs or LXC containers that can be started within minutes if the primary node fails. Traditional backups give you point‑in‑time restore but involve a restore step that adds latency. Proxmox VE’s built‑in ZFS replication scheduler creates a continuously updated copy on a remote node, giving you a warm standby with a recovery point objective (RPO) as low as the replication interval.
How Proxmox VE’s native ZFS replication works
The feature works only when the virtual disk resides on a ZFS storage pool. Each VM disk is a ZFS dataset; the scheduler takes a snapshot of the dataset, sends the incremental stream to the target node via zfs send/zfs receive, and prunes old snapshots according to a retention policy. Because the transfer is incremental after the first full send, bandwidth usage stays low and the target dataset mirrors the source at the snapshot moment.
Worked example: configuring a 15‑minute replication job for VM 101
- Ensure VM 101’s disk is on a ZFS pool (e.g.,
local-zfs). - In the web UI go to Datacenter → Replication → Add.
- Set Type to VM, VM ID to
101. - Choose the target node (e.g.,
proxmox2) and the target storage (must also be ZFS). - Define the schedule as
*/15 * * * *(every 15 minutes). - Set Prune snapshots to Keep last 4 to retain about an hour of history.
- Confirm and save; the job appears under Datacenter → Replication.
The equivalent configuration fragment that Proxmox VE writes to /etc/pve/datacenter.cfg looks like:
[replication:vm-101]
type: vm
vmid: 101
schedule: */15 * * * *
target: proxmox2
storage: local-zfs
prune-snapshots: keep-last=4
After the first run the scheduler creates a baseline snapshot, then every 15 minutes it sends only the changed blocks.
Verifying the replication and spotting issues
- Check the task log:
pvetasklogor via Node → Tasks in the UI; look for entries likeVM 101 replication job startedandfinished successfully. - On the source node list the newest snapshot:
zfs list -t snapshot -o name -s creation rpool/data/vm-101-disk-0 | tail -1 - On the target node verify that the same dataset exists and holds a snapshot with the same timestamp:
zfs list -t snapshot -o name -s creation rpool/data/vm-101-disk-0 | tail -1 - If a job fails, the UI shows a red icon and the task log contains the error (e.g.,
zfs send: dataset does not exist). Resolve by checking network connectivity, clock sync (chronydorntpd), and that the target dataset name matches the source.
Trade‑offs and limitations
- Only works with ZFS‑backed storage; VMs on LVM‑thin, directory, or iSCSI cannot be replicated this way.
- The feature’s UI and CLI options have changed between Proxmox VE 6.x, 7.x, and 8.x; verify the exact wording for your version.
- Network interruptions or prolonged pauses cause gaps in the snapshot chain; you must monitor job status and consider alerting on failed runs.
- Replication does not provide immutable storage; for long‑term off‑site retention combine it with Proxmox Backup Server (PBS) backups.
Actionable next steps
Start by identifying a ZFS‑based VM you want to protect, add a replication job with a short test interval (e.g., every 5 minutes), and verify the first snapshot appears on both nodes. Once the job runs cleanly, adjust the schedule to your desired RPO and set appropriate snapshot retention. Pair the replication with periodic PBS backups to achieve both low‑RPO warm standby and immutable archival storage.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.