Designing Harvester with Longhorn: Architecture, Trust Boundaries, and Operational Safeguards
Learn how to size a Harvester cluster with Longhorn, define trust boundaries, monitor health, and handle failure modes to keep persistent volumes reliable.
22 Mar 2026, 15:44 UTC

Problem Statement
When Harvester deploys Longhorn as the CSI‑compatible block storage layer, the cluster must guarantee that VM data survives any single node failure while keeping operational overhead low. The core question is: How should we design the cluster, define trust boundaries, and set operational checks so that persistent volumes stay reliable?
Requirements & Minimum Size
Harvester relies on the Kubernetes control plane and Longhorn’s replica mechanism. The smallest configuration that satisfies the default Longhorn replica count (3) and ensures etcd quorum is:
- Control Plane: 3 nodes, each running kube-apiserver, controller-manager, scheduler, and etcd. Minimum 4 GB RAM, 2 vCPU.
- Worker Nodes: 3 nodes to host Longhorn manager, engine, and replica pods. Each worker should have at least 8 GB RAM and 2 vCPU to leave headroom for VM workloads.
- Network: A reliable intra‑cluster network (isolated VLAN or SR‑IOV) to avoid latency spikes that can trigger replica timeouts.
With this topology, if any single worker node fails, the remaining two can still meet the 3‑replica quorum and keep volumes online.
Smallest Viable Design
Given the above, the minimal viable design is a 6‑node cluster (3 control, 3 worker). Each Longhorn replica runs as a pod on a worker node; the Longhorn manager runs as a Deployment on the control plane. The CSI driver is installed on all workers automatically by Harvester during provisioning. No additional software is required, keeping the footprint small.
Sample harvester.yaml for a 3‑worker cluster
apiVersion: harvesterhci.io/v1beta1
kind: HarvesterCluster
metadata:
name: demo
spec:
controlPlane:
replicas: 3
workerNodes:
replicas: 3
resources:
limits:
cpu: 2
memory: 8Gi
requests:
cpu: 1
memory: 4Gi
longhorn:
replicaCount: 3
nodeAffinity: true
Trust Boundaries & Data Ownership
Harvester defines two primary trust boundaries:
- Kubernetes API Server – The control plane validates all requests. Only service accounts with the
longhorn:adminrole can modify volume specs or replica counts. - Longhorn Manager Pod – Exposes a REST API that the CSI driver forwards to. VM pods interact with volumes only through the CSI driver; they never see the underlying block devices.
This separation ensures that even if a VM’s pod is compromised, it cannot tamper with the underlying storage configuration.
Operational Checks
Regular checks keep the storage layer healthy. The following tasks should be automated or monitored via alerts:
- Longhorn Health Scan – Run
longhorn health-checkevery 10 minutes to detect unhealthy replicas or engines. - Replica Rebuild Monitoring – The Longhorn UI exposes a
rebuildstatus. Usekubectl get pods -l app=longhorn-replica -o wideto confirm rebuild progress. - etcd Backup Verification – Verify that etcd snapshots exist and can be restored. Command:
etcdctl snapshot status /var/lib/etcd/snapshot.db. - Automated Snapshot Check – Example:
longhorn volume backup-list <volume-name>should return a list of snapshots. A missing snapshot indicates a scheduler failure.
Example: Checking Automated Snapshots
# Run on the Longhorn manager node or via kubectl exec
longhorn volume backup-list my-vm-volume
# Expected output contains entries like:
# 2026-10-11T21:00:00Z 2026-10-11T21:00:00Z 2026-10-11T21:00:00Z
Verify that the backup target (e.g., NFS or S3) contains the latest snapshot files.
Failure Modes & Escalation Paths
Understanding how the system behaves under stress informs design changes.
1. Node Failure
- When a worker node powers off, its replicas become unreachable.
- Longhorn automatically marks the volume as healthy with 2 replicas remaining and starts a
rebuildon a healthy node. - VMs continue to access the volume with no interruption.
2. Network Partition
- If a subset of nodes loses connectivity to the Longhorn manager, those nodes are marked unhealthy.
- Volumes on the isolated nodes switch to
read‑onlymode until the partition heals. - Automatic resume occurs once network connectivity is restored; manual intervention is rarely needed.
3. Replica Count Reduction
- Lowering
replicaCountbelow the quorum threshold (e.g., to 2) while a node fails can cause data loss during the rebalance. - Harvester warns that a
drainoperation is required before changing replicas. Failure to drain can leave the cluster in an inconsistent state.
When to Re‑Design
Certain conditions force a redesign of the cluster:
- Scaling Beyond 3 Nodes – If you add more workers, you can increase
replicaCountto 5 for higher durability, but you must also increase control plane nodes to maintain etcd quorum. - High Availability Requirements – For workloads requiring
5‑nodequorum, deploy at least 5 control plane nodes and 5 workers. - Performance‑Critical Workloads – Enable SR‑IOV for the worker nodes to reduce latency and increase throughput for Longhorn replicas.
- Disaster Recovery Needs – Add a separate backup target (S3) and enable
longhorn backup-targetto store snapshots off‑site.
Practical Validation Checklist
- Deploy a 3‑worker Harvester cluster using the ISO installer.
- Enable Longhorn via the Harvester UI.
- Create a 10 GiB volume, attach it to a test VM, and write data.
- Power‑off one worker node and confirm the VM continues to read/write.
- Run
longhorn volume backup-listto ensure snapshots exist. - Simulate a network split by disabling a NIC; observe volumes become read‑only and then resume after restoring the NIC.
- Document any alerts or manual steps required during these tests.
If all steps succeed without data loss or manual intervention, the cluster meets the design criteria.
Conclusion
By sizing the cluster to 3 control and 3 worker nodes, defining strict trust boundaries, and enforcing regular operational checks, Harvester users can rely on Longhorn to keep VM volumes available even in the face of node failures or network partitions. Adjusting replica counts or scaling the cluster should always be accompanied by a review of these design principles to avoid compromising durability.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.