Short answer: no, the Proxmox VE HA resource agent does not probe the NFS server for lock capability before starting a VM, and there is no built-in option to automatically fall back to a different locking strategy when lock acquisition fails on NFSv4. The HA stack assumes the storage layer behaves correctly; if it doesn't, the start attempt fails and the VM is left stopped. That is almost certainly what you are seeing, and it explains why the single-node test passed — with one node there is no cross-host lock contention to expose the problem.
What is confirmed vs. likely
Confirmed behavior (well-established in current PVE releases): the HA manager (pve-ha-crm/lrm) migrates the VM configuration and then issues a normal start on the target node. If the disk image cannot be opened — because of a stale lock, an idmapping mismatch, or the storage not being active/shared on that node — the start fails, the error is logged, and the CRM does not retry indefinitely. The VM stays stopped. There is no storage-capability probe and no pluggable fallback fencing mechanism for this path.
Likely explanation for your case, in order of probability:
- Stale lock or lease from the source host. If the source node did not release the image lock cleanly (crash, network partition, or forced migration), the target may refuse to open the image until the lock/lease expires.
- Inconsistent NFS mounts between nodes. Mismatched export paths, missing cache-related mount options, or NFSv4 idmapping (idmapd domain) differences can make the image appear locked or unreadable on the new host. This is easy to overlook because the mount itself looks fine.
- Storage not shared/active on the target. If the NFS storage isn't marked shared and enabled for disk images on all HA group nodes, the target is considered invalid and the VM is left stopped.
Diagnose before changing anything
The HA logs tell you exactly which case you have:
journalctl -u pve-ha-lrm -u pve-ha-crm --since "-1h"
pvesm status
qemu-img info /mnt/pve/<storage>/images/<vmid>/<disk>
Look for the start attempt and its error on the target node. A lock/permission error points at NFS; a "storage not available" error points at configuration. Confirm the image is readable from the target host with the qemu-img info check above.
Recovery
- Verify the source node is genuinely not running the VM. Do not clear locks or start the VM until you are certain — two writers on one NFS image corrupts the guest filesystem.
- Fix the underlying cause (align mount options and idmapd domains across nodes, or wait for/clear the stale lock once the source is confirmed down).
- Start the VM manually once, confirm HA reports it running, then do a controlled test migration.
Is NFSv4 unsuitable for HA?
Not inherently, but it is not "safe by default." Proxmox HA's correctness relies on watchdog-based node fencing plus storage that presents consistent locking and identity semantics on every node. The cluster cannot be made safe purely through HA settings if the NFS export behaves inconsistently — you must align the export and mount options (including idmapd domain) across all nodes. Treat that storage configuration work as a prerequisite, not an optional tuning step.
One detail that would sharpen this: the exact error line from pve-ha-lrm on the target node at migration time. If it shows a lock/permission failure, focus on NFS options and stale locks; if it shows the storage as unavailable, fix the storage definition instead.