Designing a Minimal Three‑Node Proxmox VE HA Cluster
Guide to building a minimal three‑node Proxmox VE HA cluster: requirements, design, trust boundaries, checks, and failure modes.
05 Sept 2026, 15:19 UTC

Requirements
Before building a HA cluster, verify the following prerequisites:
- At least three physical nodes running the same Proxmox VE version (8.x recommended) to achieve quorum.
- Shared storage that all nodes can access concurrently (e.g., a ZFS pool on iSCSI, Ceph RBD, or NFS).
- A dedicated, isolated network for corosync communication (private VLAN or separate NIC) to keep the trust boundary of the HA ring.
Smallest suitable design
Node software
Each node runs Proxmox VE 8.x with the standard pve-cluster, corosync, and pacemaker packages. The HA resources are managed by the ha-manager service.
Fence agent
A simple fence agent such as fence_ipmilan (IPMI) or fence_manual for testing is configured on every node. The agent must be able to power‑off a peer node when corosync detects a loss of quorum.
# Example fence_ipmilan entry in /etc/pve/ha/resources.cfg (run as root on any node)
# Replace , , with actual values
ipmi: login= passwd= lanplus=1
Place this file under /etc/pve/ha/resources.cfg and reload the HA manager: systemctl reload ha-manager.
Trust and data boundaries
- The corosync ring is trusted only on the isolated HA network; storage traffic may flow over a separate trusted storage network.
- VM/CT disks reside on the shared storage, so they never traverse the HA ring.
- The Proxmox web GUI and API are protected by TLS and role‑based access control (RBAC), independent of the HA subnet.
Operational checks
- Check cluster health:
pvecm statuson each node. Look forQuorum: yesand all nodes in the member list. - Verify fence agent logs:
journalctl -u pve-fencedor/var/log/daemon.logfor successful fence commands. - Test resource migration:
ha-manager set --state startedon a node, then power off that node and confirm the VM starts on another node viaha-manager status. - Validate shared storage accessibility after a node failure: for ZFS, run
zpool status -von the surviving nodes; for Ceph, runceph statusand ensureHEALTH_OK.
Failure modes and design‑change triggers
Split‑brain due to corosync ring loss
If the HA network partition prevents a majority of nodes from communicating, the cluster loses quorum. Recovery requires manual fencing of the isolated partition or deploying a quorum device (e.g., a separate net‑device or a third‑party quorum server).
Storage loss
When the shared storage becomes unavailable, VMs/CTs encounter I/O errors. The HA manager cannot migrate running workloads because the disks are inaccessible. In this case, consider adding a replicated storage layer (Ceph with erasure coding or ZFS send/receive to a secondary site) or switching to a different storage backend.
Fence agent insufficiency
If IPMI or the chosen fence method fails (e.g., loss of BMC connectivity), the cluster cannot power‑off a stuck node. Adding a watchdog that triggers a hard reset, or configuring a redundant fence method (dual IPMI channels or a smart PDU), changes the design to improve reliability.
Regularly update Proxmox VE and the corosync/pacemaker packages; older versions may lack fixes for known quorum or fencing bugs that could precipitate split‑brain scenarios.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.