Automating VM Failover with Proxmox VE High Availability
Learn how Proxmox VE HA automatically restarts VMs after a host failure, reducing downtime to seconds with shared storage and fencing.
24 Dec 2025, 07:07 UTC

Problem: a single host failure takes down everything
When a Proxmox VE node loses power or crashes, all virtual machines and containers running on that host become unavailable until an administrator manually moves them to another node. This downtime can stretch from minutes to hours, especially in small clusters where spare capacity is limited.
Thesis: HA with shared storage and fencing automates recovery
Enabling Proxmox VE High Availability (HA) adds a cluster‑wide manager that monitors resources, detects node loss via Corosync/Pacemaker, fences the failed host (using IPMI, iLO, or similar), and restarts the affected VM/CT on a healthy node. With reliable fencing and shared storage, failover typically completes in 10‑15 seconds, turning a manual scramble into an automated safeguard.
HA Architecture Basics
HA requires three core pieces:
- Cluster of at least three nodes (to maintain quorum).
- Shared storage accessible by every node (Ceph RBD, NFS, or iSCSI).
- HA manager service (pve-ha-manager) that works with Corosync/Pacemaker to monitor resources and trigger fencing/migration.
The cluster uses a voting mechanism: if a node disappears, the remaining nodes retain quorum, declare the missing node lost, and instruct the fencing agent to power it off (if not already off). Once fencing succeeds, the HA manager relocates any VM/CT marked as HA‑managed to another node with sufficient resources.
Enabling HA on a Three‑Node Cluster
- Install the required packages on every node (run as root):
apt update && apt install -y corosync pve-ha-manager pve-cluster - Create or join the cluster (example from node1):
pvecm create cluster01 # on first node pvecm add node1_ip # on node2 and node3 - Configure shared storage in the datacenter (e.g., an NFS export mounted at /mnt/pve-ha-storage on all nodes) and add it as a storage entry in the GUI or via
pvesm add nfs ha-storage --server nfs-server --export /path --mntpoint /mnt/pve-ha-storage --content images,rootdir. - Edit the HA resources file to define which VMs/CTs are managed:
Replace the IDs with your own VM/CT numbers.# /etc/pve/ha/resources.cfg vm:100 ct:200 - Start the HA service:
systemctl enable --now pve-ha-manager - Verify cluster health and HA status:
pvecm status # should show quorum and 3 nodes ha-manager status # lists resources and their current node
Worked Failover Example
To see HA in action, simulate a host failure on node2:
- Ensure IPMI fencing is configured for node2 in
/etc/pve/ha/fencing.cfg(example):node2: ipmi addr=192.168.1.12 user=admin password=secret - Power off node2 (run as root on that node):
systemctl poweroffRisk: If fencing fails, the cluster may attempt to start the VM on another node while the original host is still powered, risking disk corruption. Verify fencing works before relying on HA.
- Observe the HA manager logs on any surviving node:
You should see entries similar to:tail -f /var/log/ha/ha.log- "node2 lost"
- "fencing node2 via IPMI succeeded"
- "migrating VM 100 to node3"
- "VM 100 started on node3"
- Confirm the VM is running on the new node:
ha-manager status | grep vm-100 qm list | grep 100
The entire process typically finishes within 10‑15 seconds, after which services resume automatically.
Trade‑offs and Practical Checks
- Fencing reliability: HA depends on the fencing agent reaching the host’s management controller. Test fencing manually (
ha-manager debug) before putting VMs into production. - Network partitions: A split‑brain scenario can cause duplicate VM starts if quorum is lost. Use a redundant network and consider a quorum device.
- Storage latency: Shared storage adds latency; monitor I/O wait times and ensure the storage subsystem can handle the combined load of all nodes.
- Resource overhead: The HA manager consumes minimal CPU/RAM, but each failover triggers a full VM migration, which temporarily uses network bandwidth.
Regular verification steps:
- Run
pvecm statusdaily to confirm quorum and node count. - Check
ha-manager statusfor any resources showing "stopped" or "failed". - Review
/var/log/ha/ha.logfor fencing or migration errors after any maintenance.
Closing advice
Start with a two‑node test cluster to validate fencing and shared storage, then add a third node to achieve quorum. Keep the HA manager logs under surveillance, and document your fencing credentials securely. When the architecture is sound, Proxmox VE HA turns an unpredictable host outage into a brief, automated blip—keeping your services running with minimal manual intervention.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.