Harvester and Longhorn: Why Replication Factor Is Your First Storage Decision
Harvester uses Longhorn as its default storage engine to provide distributed block storage for VMs. The replication factor you choose drives durability, storage overhead and write latency, and the storage plane network determines whether that durability actually holds.
15 Jul 2026, 23:36 UTC

You have a three-node Harvester cluster and a VM that must survive a host failure without an external SAN. The useful takeaway is that Harvester does not give you local disks with a network share. It gives you Kubernetes-native block volumes managed by Longhorn, and the durability you get is a direct function of the replication factor you choose and the storage plane network you provide.
The problem local disks create in HCI
Hyper-Converged Infrastructure, HCI, means compute and storage live on the same nodes. With plain local disks a VM disk is tied to the node it lives on. A host reboot or hardware fault makes the disk unavailable until the VM is rescheduled. Harvester avoids that single point of failure by treating VM disks as persistent volumes provisioned through a Container Storage Interface, CSI, driver. CSI is the Kubernetes standard for plugging storage into workloads. In Harvester, the driver is Longhorn.
Longhorn is Harvester’s default storage engine. It uses a shared-nothing architecture: each node contributes local disks, and Longhorn replicates data blocks across nodes. There is no central storage appliance. The VM disk is a Longhorn volume, and the Harvester virtualization layer attaches it via CSI.
How replication and access modes work in practice
Shared-nothing replication, not a SAN
A Longhorn volume is a set of replicas. A replica is a full copy of the volume data stored on a node’s designated storage path. Writes are synchronously replicated to a quorum of replicas before acknowledgment. That design removes a single point of failure, but it also means write latency and storage overhead scale with replication factor.
Read-Write-Many, RWX, is not native to block volumes. Harvester enables shared filesystems through Longhorn’s integrated NFS server. Longhorn can export a volume as NFS for workloads that need concurrent file access. That is a separate access mode from the block volumes used for VM disks.
Snapshots and backups as recovery primitives
Longhorn supports point-in-time snapshots of volumes and scheduled backups to an offsite target. Snapshots are volume-consistent and can be used for cloning or rollback. Backups enable offsite data migration and disaster recovery. Both are managed through the Longhorn UI and the Harvester UI.
A concrete decision: replication factor for a VM disk
Replication factor is the engineering decision that matters most on day one. The default is often 3. A factor of 3 means three replicas on three different nodes. A factor of 2 reduces overhead but tolerates fewer failures.
Example decision context. You create a VM in Harvester with a 100 GiB disk. With replication factor 3, the cluster must reserve roughly 300 GiB of raw capacity across nodes, minus overhead. With replication factor 2, the reservation is roughly 200 GiB. The choice is not just capacity. Synchronous replication means each write must be acknowledged by the replica set over the storage network. Higher replication and higher network latency increase write latency for the VM.
Where to inspect this. In the Harvester UI, open the VM, view the attached disks, and open the corresponding Longhorn volume. The Longhorn UI shows replica distribution and synchronization status per volume. You can also query from a machine with cluster access and appropriate RBAC, for example a workstation with the Harvester kubeconfig.
kubectl -n longhorn-system get longhornvolume <volume-name>Run this with a user that has read access to the longhorn-system namespace. A meaningful check is that the number of healthy replicas matches the desired replication factor and that replicas are spread across nodes. Do not modify replica count on a live production volume without a maintenance window; changing replication triggers data rebalancing which consumes network and I/O.
Trade-offs and limitations you will feel
High replication factors increase storage overhead linearly and can impact write latency due to synchronous network replication. Performance is heavily dependent on underlying physical disk speed and network throughput between nodes. A 10 GbE storage plane with NVMe nodes behaves very differently from 1 GbE with SATA.
Incorrect network configuration for the storage plane can lead to split-brain scenarios or volume detachment. Split-brain means two replicas think they are primary. Harvester relies on a correctly isolated, low-latency network for Longhorn traffic. Verify that the storage network is separate from management and VM traffic, and that firewall rules allow Longhorn ports between nodes.
Longhorn NFS for RWX adds another operational surface. It is useful for shared filesystems, but it is not a replacement for block replication for VM disks.
How to verify the setup is actually resilient
A practical verification is to observe replica placement and then confirm VM accessibility after a controlled node change. Inspect the Longhorn UI to confirm each volume has the expected number of replicas on distinct nodes and that the replicas report healthy and in sync.
For a resilience check, deploy a test VM on a three-node cluster with a volume using replication factor 3. With the cluster under no production load, observe the VM remains accessible after a planned node drain. This is a verification step, not a production procedure. Risks include data unavailability if the storage network is misconfigured or if too many nodes are down at once.
Close by treating replication factor as a capacity and latency budget, not just a durability toggle. Document your storage network topology, keep Longhorn replicas balanced across nodes, and monitor replica health as part of routine operations.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.