OpenStack Nova Live Migration: Minimal Architecture Note
A concise architecture note for OpenStack Nova live migration: requirements, minimal two‑node design, trust boundaries, pre‑migration checks, typical failure modes, and conditions that force a design change.
24 Jun 2026, 17:13 UTC

Requirements
To enable Nova live migration you need:
- Shared storage that all compute nodes can read and write with low latency (e.g., NFS export, Ceph RBD pool).
- Libvirt/QEMU-KVM version that supports live block copy (virsh >= 1.0.0).
- A functional control plane: nova-conductor, nova-scheduler, and nova-compute services running on each host.
- Neutron configured so that the same VLAN/VXLAN network segment is present on both source and destination computes.
Smallest Suitable Design
A two‑node Nova deployment satisfies the above:
- Node A and Node B each run nova‑compute.
- Both nodes mount the same NFS export at
/var/lib/nova/instances(or use a common Ceph pool). - Neutron provides a single provider network (e.g., VLAN 101) that is plugged into both computes.
- The scheduler places instances on whichever node has free RAM and CPU.
This is the minimal topology; adding more nodes only scales the same pattern.
Trust/Data Boundaries
The migration stream consists of memory pages and, if live block copy is used, disk data. By default libvirt sends this data over the private compute‑node network without encryption. Therefore the trust boundary is:
- The internal network linking the two compute nodes (typically a dedicated migration VLAN).
- The shared storage backend, which must be trusted to hold the instance’s disk images.
If you enable TLS in libvirt (migration_tls=1 in libvirtd.conf), the boundary shifts to the TLS termination points on each compute node.
Operational Checks
Before issuing a live‑migration request, verify:
- CPU feature compatibility: the source and destination must expose the same CPU model (or use CPU mode masking via
cpu_mode=customandcpu_modelin nova.conf). - Sufficient free RAM on the target host (at least the instance’s defined memory).
- Enough free disk space on the shared storage for any temporary migration files.
- That the instance does not use hardware‑specific devices that block live migration (PCI passthrough, SR‑IOV VF, GPU passthrough, USB devices).
- That the libvirt migration port range (default 49152‑49215) is open in both directions between the computes (check with
nc -z <dest> 49152-49215).
Failure Modes
- Network interruption: If the migration VLAN drops packets, libvirt aborts and the instance remains running on the source host; the destination side cleans up any partial memory state.
- Storage latency spikes: High latency on the NFS/Ceph backend can cause the live block copy phase to exceed the timeout (
live_migration_timeoutin nova.conf), leading to a failed migration and a revert to the source. - Libvirt version mismatch: Different libvirt releases may interpret the migration protocol differently, resulting in silent failures; check
virsh versionon both nodes and keep them within the same minor release. - Incompatible CPU models: Without CPU mode masking, a guest may encounter a missing instruction and crash during the resume phase on the destination.
Design‑Change Conditions
The architecture must be revisited when any of the following occurs:
- Local‑only storage is introduced (e.g., each compute uses its own disk). Live migration cannot work because the disk image cannot be accessed concurrently; you must switch to cold‑migration or storage‑live‑migration (if supported).
- The workload requires SR‑IOV, GPU passthrough, or other PCI devices that libvirt marks as non‑migratable. In this case disable live migration for those flavors or use a workload‑specific host aggregate.
- You enable encrypted migration streams (TLS). This adds a trust boundary at the libvirt TLS endpoints and requires certificate management; the migration network can then be treated as untrusted.
- You change the shared storage technology to one with higher latency (e.g., moving from Ceph RBD over 10 GbE to NFS over a 1 GbE link). You may need to increase
live_migration_timeoutor upgrade the storage network to avoid timeouts.
Practical Verification (Example)
Deploy a minimal DevStack with two compute nodes:
# On both nodes, add to local.conf: [[local|localrc]] ENABLED_SERVICES+=n-cpu,n-cond,n-sch # Shared NFS export (assume server at 10.0.0.100:/export/nova) NFS_SERVER=10.0.0.100 NFS_EXPORT=/export/nova NFS_MOUNT_POINT=/var/lib/nova/instances # Enable live migration flag LIVE_MIGRATION_FLAG=VIR_MIGRATE_LIVE
After stacking, create a test instance:
# Source node (where nova‑compute is active) openstack server create --flavor m1.tiny --image cirrus-0.5.2 --nic net-id=$(openstack network list -c ID -f value) test-vm
Run the migration (replace compute2 with the host name of the target nova‑compute service):
openstack server migrate --live compute2 test-vm
Monitor progress:
openstack server show test-vm -c status -c OS-EXT-STS:vm_state -f value
The instance should stay reachable (ping/SSH) throughout. To test a failure condition, add latency between the nodes:
# On the source node, add 100ms delay to the migration VLAN (replace eth1 with your migration interface) sudo tc qdisc add dev eth1 root netem delay 100ms
Repeat the migration command; observe that it either takes longer than the default timeout or aborts, and check /var/log/nova/nova-compute.log for messages like “migration timed out”.
To demonstrate the hardware boundary, attach a PCI passthrough device to the instance before migration and repeat the command; the migration will fail with a libvirt error similar to “operation not supported: device is not migratable”.
Limitations
This note assumes a homogeneous CPU architecture across hosts and a shared storage backend that provides block‑level access. It does not cover storage‑live‑migration (where the disk image moves while the instance runs) or multi‑cell Nova deployments. Verify any production‑specific tuning (e.g., live_migration_bandwidth, live_migration_downtime) against your workload’s SLA.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.