vSphere HA, vMotion, and Fault Tolerance: Architecture for Zero Downtime
Deploy VMware vSphere HA with vMotion and Fault Tolerance for zero downtime on critical workloads. This architecture guide covers minimal design, trust boundaries, operational checks, and scaling considerations.
15 Aug 2025, 04:05 UTC

Problem Statement
Critical virtual machines must stay online even when a host, storage, or network component fails. VMware vSphere provides three mechanisms to achieve this: High Availability (HA) restarts VMs on healthy hosts, vMotion enables live migration without downtime, and Fault Tolerance (FT) creates a live shadow VM for continuous availability. The challenge is designing a minimal cluster that guarantees zero downtime for mission-critical workloads while managing resource constraints.
Requirements Checklist
- vSphere 7.x or later with vCenter Server; Enterprise Plus license required for FT
- Minimum two ESXi hosts in a single cluster
- Shared storage (vSAN, NFS, or SAN) accessible by all hosts
- Network latency below 5 ms for FT; 10 Gbps bandwidth recommended for vMotion
- CPU virtualization extensions enabled (Intel VT-x/EPT or AMD-V/RVI)
- VM vCPU count limited to 2x physical cores when FT is enabled
Smallest Suitable Design
A 2-host cluster can provide zero downtime for critical workloads with the following configuration:
- Cluster Configuration
- Enable DRS in Manual mode to control VM placement
- Enable HA with Admission Control set to 1 host failure tolerance
- Enable vMotion and FT for all hosts in the cluster
- Storage
- Single datastore (vSAN or NFS) shared by all hosts
- Ensure sufficient IOPS and throughput for VM workloads
- Store FT shadow VMs on the same datastore
- Network
- Dedicated VMkernel port group for vMotion traffic
- Separate VMkernel port group for FT traffic if required
- Consider vSphere Replication for cross-site protection
- Fault Tolerance
- Enable FT on critical VMs with compatible CPUs
- Monitor FT health state via vSphere Client or PowerCLI
Trust and Data Boundaries
- All VMs share the same storage backend; data consistency depends on storage controller reliability
- vMotion traffic should be isolated on dedicated VMkernel to prevent congestion
- FT shadow VMs are separate inventory objects but share host resources; enforce strict access controls
- Use vCenter RBAC to limit HA and FT management to authorized operations personnel
Operational Checks
Regular validation ensures the design remains healthy. The following table summarizes critical checks, tools, and expected results.
| Check | Tool/Command | Expected Result |
|---|---|---|
| HA Test-Failover | vSphere Client → Cluster → Configure → vSphere HA → Test Fail-over | All VMs restart on surviving host within grace period |
| FT Health State | Get-FTHealth -VM <VM_NAME> | State = Active, no latency warnings |
| Network Latency | Test-Connection -ComputerName <HOST> -Count 10 or vSphere Network I/O | Latency below 5 ms for FT traffic |
| CPU Utilization | vCenter Performance charts → CPU % | CPU usage below 80% normally, below 60% during simulated failure |
| Storage I/O | vCenter Storage performance charts → IOPS | IOPS below 80% of rated capacity for FT VMs |
Example PowerCLI Script
Run this script from the vCenter PowerCLI console with Administrator privileges to enable HA, vMotion, and FT on a cluster.
# Connect to vCenter
Connect-VIServer -Server <vcenter> -User <user> -Password <password>
# Enable HA, vMotion, FT on cluster
$cluster = Get-Cluster -Name <CLUSTER_NAME>
$cluster | Set-Cluster -HAEnabled $true -VMotionEnabled $true -FTEnabled $true
# Verify FT health for a VM
$vm = Get-VM -Name <VM_NAME>
Get-FTHealth -VM $vm
Failure Modes and Mitigation
- Host Failure
- HA restarts VMs on remaining host within configured grace period
- Mitigation: Maintain admission control for 1 host failure; monitor CPU overhead
- Storage Failure
- Both HA and FT depend on shared storage availability
- Mitigation: Use vSAN with redundancy or multi-path SAN with backup controller
- Network Partition
- FT requires low-latency, high-bandwidth links for shadow VM synchronization
- Mitigation: Deploy redundant uplinks with NIC teaming and load balancing
- CPU Saturation
- FT doubles CPU usage; small clusters may become over-committed
- Mitigation: Ensure adequate CPU headroom; separate non-critical VMs to different cluster
When to Change the Design
- Scaling Out – Add a third host and enable DRS in Fully Automated mode for load balancing
- High I/O Workloads – Replace NFS with dedicated SAN or upgrade vSAN tier
- Compliance Requirements – Isolate FT shadow VMs on dedicated datastore with separate access controls
- Budget Constraints – Drop FT and rely on HA with scheduled backups and fail-over tests
- CPU Compatibility Issues – Upgrade CPUs or create separate cluster for FT-protected VMs
Practical Verification Checklist
- Run HA test-failover at least weekly
- Verify FT health state daily via vSphere Client or automated script
- Monitor CPU and storage metrics continuously with alerts for >80% CPU or >90% IOPS
- Document any design deviations and update architecture documentation
Limitations
- FT consumes approximately 100% additional CPU; small clusters risk over-commitment
- Storage performance directly impacts HA restart times; suboptimal I/O extends downtime
- FT requires Enterprise Plus license and compatible CPU virtualization extensions
- vMotion requires consistent VMkernel port group configuration across all cluster hosts
Conclusion
Combining HA, vMotion, and FT in a minimal 2-host cluster delivers near-zero downtime for critical workloads. Success depends on starting with a simple, well-defined design, maintaining strict trust boundaries, and performing regular operational checks. As requirements evolve, adapt the architecture by scaling resources, upgrading components, or modifying protection strategies.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.