Live Migrating OpenStack Nova Instances Between Compute Hosts Using Shared Storage
Step-by-step guide for live migrating running OpenStack Nova instances between compute hosts using shared storage (Ceph, NFS, iSCSI). Covers prerequisites, migration commands, monitoring, verification, and recovery from common failures.
22 Jun 2026, 22:57 UTC

The Problem: Moving Running Workloads Without Downtime
When a compute host needs maintenance, rebalancing, or replacement, you must move running instances elsewhere. Cold migration powers off the VM, causing unacceptable downtime for production workloads. Nova live migration keeps the instance running while transferring its memory and state to another host, provided both hosts share access to the instance's disk images.
Desired Outcome
A running instance moves from a source compute host to a target compute host with minimal service interruption. The instance retains its IP addresses, memory state, and network connections. After migration, the instance runs on the target host and the source host shows the instance as migrated.
Prerequisites
- Shared storage backend: Ceph RBD, NFS, or iSCSI accessible from both source and target hosts. The instance's root disk and any attached volumes must reside on this shared storage. Local disk (ephemeral) instances cannot use pre-copy live migration.
- Compatible CPU architectures: Source and target hosts must have compatible CPU models. Migration between x86_64 and ARM fails unless cross-architecture support is explicitly configured (rare and not recommended for production).
- Libvirt and QEMU versions aligned: Major version mismatches can cause migration failures. Run
virsh versionon both hosts to confirm compatibility. - Network connectivity: Dedicated migration network (typically 10 Gbps or higher) between compute hosts. The
live_migration_inbound_addrinnova.confshould point to this network interface. - Nova compute services healthy: Both source and target hosts report
enabledandupstatus.
Verify Compute Host Status
Run on the controller node (or any host with admin OpenStack RC file sourced):
openstack compute service list --service nova-compute
Confirm both source and target hosts show Status: enabled and State: up. If a host shows down, investigate nova-compute service logs before proceeding.
Identify the Instance and Target Host
List instances on the source host:
openstack server list --host compute-01 --all-projects
Note the instance UUID (e.g., a1b2c3d4-e5f6-7890-abcd-ef1234567890). Choose a target host that passes scheduler filters. To see eligible hosts:
openstack hypervisor list
For manual host selection, use the target hostname (e.g., compute-02).
Initiate Live Migration
Run from the controller or any host with admin credentials:
openstack server migrate --live compute-02 \
--wait a1b2c3d4-e5f6-7890-abcd-ef1234567890
Flags explained:
--live: Use live migration (pre-copy with shared storage). Without this flag, Nova performs cold migration.--live compute-02: Target hostname. Omit to let the scheduler choose.--wait: Block until migration completes or fails, showing progress.
Permissions: Requires admin role or os_compute_api:os-migrate-server:migrate_live policy permission.
Monitor Migration Progress
The --wait flag shows stages: preparing, running, post-copy (if enabled), completed. For detailed libvirt-level visibility, tail the libvirt log on the source host:
tail -f /var/log/libvirt/libvirtd.log | grep -i migrate
Expected log sequence: migration started → iteration N (memory pages sync) → migration completed.
Verify Migration Completion
After the command returns success, confirm the instance now runs on the target host:
openstack server show a1b2c3d4-e5f6-7890-abcd-ef1234567890 \
-f value -c OS-EXT-SRV-ATTR:host
Output should show compute-02. Also verify instance status:
openstack server show a1b2c3d4-e5f6-7890-abcd-ef1234567890 \
-f value -c status
Should return ACTIVE.
Check Instance Health Post-Migration
SSH into the instance (if networking permits) and verify:
- Application processes running
- Network connectivity intact
- Disk I/O normal (
iostat -x 1) - No kernel messages indicating hardware changes (
dmesg -T)
For Windows instances, check Event Viewer for unexpected device changes.
Common Failure Modes and Recovery
Migration Times Out or Stalls
Large memory instances (64 GB+) with high page dirty rates may not converge. The pre-copy iterations continue until remaining dirty pages fit in the final cutover window. If the instance writes memory faster than the network transfers, migration stalls.
Recovery: Abort with openstack server migrate --abort <uuid> (run on controller). The instance continues running on the source host. Options:
- Schedule migration during lower write activity
- Increase
live_migration_completion_timeoutinnova.conf(default 800 seconds) - Enable post-copy fallback: set
live_migration_permit_post_copy = trueandlive_migration_permit_auto_converge = trueinnova.confon both hosts, then restartnova-compute
Network Interruption During Migration
If the migration network fails mid-transfer, the instance may crash or enter an error state.
Recovery: Check instance status. If ERROR, try openstack server reset-state --active <uuid> then hard reboot: openstack server reboot --hard <uuid>. This restarts the instance on the source host (if migration hadn't committed) or target host (if cutover completed).
Target Host Resource Exhaustion
Target host runs out of memory or CPU during cutover.
Recovery: Migration fails automatically. Instance remains on source. Check nova-compute logs on target for MigrationPreCopyFailed or InsufficientResources. Free resources or select a different target.
Limitations
- No shared storage = no pre-copy: Without Ceph/NFS/iSCSI, only post-copy or shared-nothing migration works, which carries higher risk and downtime.
- PCI/GPU passthrough devices: Instances with PCI passthrough (GPUs, NICs) cannot live migrate unless the target host has identical devices and
pci_passthrough_whitelistconfigured. - NUMA topology: Instances with explicit NUMA pinning require target host with matching topology.
- Downtime not zero: Final cutover pauses the VM for typically 100 ms–2 s. For latency-sensitive workloads, test first.
Practical Verification Checklist
- Run
openstack compute service list— both hostsenabled/up - Confirm instance on shared storage:
openstack volume list --server <uuid>shows volumes on Ceph/NFS backend - Execute live migration with
--wait - Verify
OS-EXT-SRV-ATTR:hostchanged to target - Confirm instance
ACTIVEand application responsive - Check
/var/log/libvirt/libvirtd.logon source formigration completedwith exit code 0
Rollback Consideration
Live migration is a state-changing operation. Once cutover completes, the instance runs on the target host. There is no automatic rollback. To move the instance back, run another live migration targeting the original host. Plan maintenance windows accordingly.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.