Putting a Harvester Node into Maintenance Mode: Automatic VM Evacuation and Safe Hardware Service
How to put a Harvester node into maintenance mode so VMs live-migrate automatically, what to check before powering hardware off, and how to recover from stalled evacuations.
28 Oct 2025, 05:48 UTC

Desired Outcome
Place a Harvester node into maintenance mode so its running virtual machines are automatically live-migrated to other nodes. The cluster stays available while you safely service or replace hardware on the evacuated node.
Prerequisites
- A Harvester cluster (v1.2 or later; behavior described here assumes a current 1.x release) with at least three healthy nodes, so quorum and capacity survive losing one node.
- VMs using shared, network-attached storage (Longhorn volumes, the default in Harvester) rather than node-local disks.
- Sufficient CPU and memory headroom on the remaining nodes to absorb the evacuated VMs.
- No VMs on the target node using PCI/GPU passthrough, USB host devices, or other host-pinned resources – these cannot be live-migrated.
- Access to the Harvester UI with permissions to manage nodes, or
kubectlaccess to the cluster.
Procedure
- Check cluster health first. In the Harvester UI, confirm all other nodes show
Readyand that storage health is green. Do not start maintenance while another node is degraded. - Enable maintenance mode. In the UI go to
Hosts, select the target node, and chooseEnable Maintenance Modefrom the node actions menu. Harvester cordons the node and begins evicting workloads; KubeVirt live-migrates each migratable VirtualMachineInstance (VMI) to another node. - Monitor evacuation. On the node detail page, watch the VM count drop toward zero. From a management workstation with cluster access you can also watch VMIs:
kubectl get vmis -A -o wideEach migrated VMI should show a new node in the
NODEcolumn and remain in theRunningphase. Non-migratable VMs (passthrough devices, local disks) will block the drain and must be shut down manually before the node finishes entering maintenance. - Confirm the node is in maintenance. The node status in the UI changes to
Maintenance, and only system pods remain. Verify no user VMs are left before touching hardware. - Perform the hardware work. Shut down or power off the node via its BMC/IPMI once evacuation is complete.
- Return the node to service. Power the node back on and let it rejoin the cluster. Then in the UI select the node and choose
Disable Maintenance Mode(this uncordons it). VMs are not automatically moved back; they rebalance only when restarted or migrated, which is normal.
Expected Checks
- Node shows
Maintenancestatus and zero running VMs before power-off. - Every migrated VM shows
Runningwith a new node assignment; spot-check guest connectivity (ping, SSH, or application health check) from inside or outside the guest. - Longhorn volume health stays
HealthyorDegraded-but-rebuildingduring the operation; investigate any volume stuck in fault before proceeding. - After re-enabling, the node returns to
Readyand passes the UI health checks.
Recovery Options
- Stalled migration: If a VMI will not migrate (commonly due to passthrough devices or a migration timeout), shut that VM down gracefully from the guest or UI, let the drain finish, and start it again on another node after maintenance – or before, if capacity allows.
- Abort maintenance: Select
Disable Maintenance Modeto uncordon the node. Already-migrated VMs stay where they are; remaining VMs continue running on the node. - Insufficient capacity: If target nodes lack headroom, evacuations pend. Free resources (stop non-critical VMs) or add a node, then retry.
Limitations
- Live migration requires shared storage and no host-pinned devices; plan manual shutdowns for exceptions.
- Maintenance-mode behavior, migration timeouts, and UI options vary between Harvester releases – validate the exact steps against your deployed version's documentation before relying on them in production.
Practical Verification
After the full cycle, confirm the node is schedulable again:
kubectl get nodes
The serviced node should show Ready with no SchedulingDisabled marker. Migrate or restart one test VM onto it to prove it can host workloads normally.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.