Using Harvester’s Longhorn Integration for Hyper‑Converged Infrastructure: A Practical Guide
Learn how Harvester’s Longhorn‑backed block storage makes hyper‑converged infrastructure resilient, the trade‑offs of synchronous replication, and a step‑by‑step example to create a replicated volume and recover a VM after node failure.
19 Jun 2026, 00:05 UTC

Problem & Thesis
Hyper‑converged infrastructure (HCI) bundles compute, storage, and networking into a single platform. A common pain point is ensuring that virtual machine (VM) data survives node outages without manual intervention. Harvester, a lightweight Kubernetes‑based HCI stack, solves this by integrating Longhorn, a cloud‑native distributed block storage system. The key benefit: a VM can be live‑migrated or restarted on any healthy node while its data remains intact. This blog walks through how to create a replicated Longhorn volume, attach it to a VM, simulate a node failure, and verify recovery—highlighting synchronous replication’s latency trade‑off and practical monitoring tips.
Architecture Overview
Harvester runs on a cluster of Linux nodes. Each node exposes a kubelet and a Longhorn engine container. Longhorn uses Custom Resource Definitions (CRDs) to model Volume objects. When a volume is defined with a replicaCount, Longhorn creates that many replicas on distinct nodes. Replicas are kept in sync via synchronous replication: every write is acknowledged on all replicas before the VM sees it. This guarantees data durability at the cost of write latency proportional to the slowest replica.
Key Components
- Harvester Controller – orchestrates VM lifecycle and volume attachment.
- Longhorn Engine – runs on each node to provide the block device.
- Replica Pods – one per volume per node, storing the actual data.
- Volume CRD – declarative spec for size, replica count, and access mode.
Creating a Replicated Volume
Below is a minimal YAML that creates a 10 GiB volume with three replicas. Run it on any node that has kubectl configured for the Harvester cluster.
apiVersion: longhorn.io/v1beta1
kind: Volume
metadata:
name: vm-data
spec:
size: 10Gi
numberOfReplicas: 3
numberOfScheduledNodes: 3
# Optional: specify nodes explicitly
# nodeSelector:
# matchLabels:
# kubernetes.io/hostname: <node-name>
Apply the manifest:
kubectl apply -f vm-data.yaml
Check status:
kubectl get volumes.longhorn.io vm-data -o wide
Expected output shows State: Attached and ReplicaCount: 3. Verify that each replica is on a different node:
kubectl get replicas.vm-data -o jsonpath='{.items[*].nodeId}'
Risks: if the cluster has fewer than three nodes, the replica creation will stall. Ensure enough capacity before scaling.
Attaching the Volume to a VM
Harvester exposes a harvesterhci.io/volume annotation on VM specs. Example VM YAML:
apiVersion: kubevirt.io/v1alpha3
kind: VirtualMachine
metadata:
name: demo-vm
spec:
template:
spec:
domain:
devices:
disks:
- name: disk0
disk: {}
bus: virtio
volumes:
- name: disk0
persistentVolumeClaim:
claimName: vm-data
Apply and watch the VM come online. The volume is presented as a block device to the guest OS.
Simulating a Node Failure
To test resilience, cordon and delete a node that hosts one of the replicas:
kubectl cordon <node-name>
kubectl delete node <node-name>
Longhorn automatically marks the replica on the failed node as Unavailable and promotes a new replica to maintain the desired count. Verify:
kubectl get replicas.vm-data -o wide
Now restart the VM (or wait for Harvester to reschedule it). The VM boots onto a healthy node, and the block device remains intact because the data was replicated synchronously before the failure.
Trade‑offs & Limitations
- Write Latency – synchronous replication means each write must reach all replicas. In a 3‑node cluster over a 1 Gbps link, latency can add ~1–2 ms per I/O, noticeable under heavy workloads.
- Resource Overhead – each replica consumes CPU, memory, and disk space. A 10 GiB volume with 3 replicas uses ~30 GiB of raw storage plus overhead.
- Network Dependence – replication traffic traverses the overlay network. Poor network health can cause replication lag and eventual consistency issues.
- Node Homogeneity – Harvester assumes all nodes provide similar performance. Mixing high‑end and low‑end nodes can skew replica placement and affect I/O.
Practical check: monitor replication lag with kubectl get replicas.vm-data -o jsonpath='{.items[*].replicaState}' and network I/O with sar -n DEV 1 10 on each node. If lag spikes, consider increasing replicaCount or upgrading network bandwidth.
Actionable Takeaways
- Create a 10 GiB volume with
numberOfReplicas: 3to guarantee data durability. - Attach the volume to VMs using a PVC reference; Harvester handles block device exposure.
- Test node failure by cordoning/deleting a node and verifying VM recovery.
- Monitor write latency and replication lag; adjust
replicaCountor network capacity as needed. - Document your cluster’s node count and network topology to avoid replication stalls during scaling.
By following this workflow, you can confidently deploy hyper‑converged workloads on Harvester, knowing that Longhorn’s synchronous replication keeps your VM data safe even when individual compute nodes go down.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.