Choosing K3s HA: When dqlite Beats External etcd
Stop managing separate etcd clusters for small K8s deployments. Learn how K3s uses embedded dqlite to provide high availability with a tiny memory footprint and simplified operations.
13 Jan 2026, 18:32 UTC

The High Availability Dilemma in Edge Clusters
Setting up a High Availability (HA) Kubernetes control plane usually requires a trade-off: you either manage a complex external etcd cluster—which demands its own certificates, dedicated ports, and significant RAM—or you accept a single point of failure. For edge deployments or small-to-medium clusters, the overhead of etcd often outweighs the benefits of the workloads it manages.
The takeaway is simple: K3s solves this by embedding dqlite (distributed SQLite). This allows you to run a fully redundant, Raft-based control plane across three or more nodes without installing any external database software. You get the linearizable consistency of etcd with the footprint of a local file.
How dqlite Simplifies the Control Plane
Unlike standard Kubernetes, which treats etcd as a separate process or cluster, K3s integrates dqlite directly into the server process. dqlite uses the Raft consensus algorithm to ensure that all server nodes agree on the state of the cluster, but it stores that state in a SQLite database file on disk.
This architectural shift removes several operational hurdles:
- Resource Efficiency: While etcd can easily consume several gigabytes of RAM, dqlite typically keeps the state machine in a small file (often 10–50 MB), significantly lowering the memory floor for control-plane nodes.
- Network Simplification: You no longer need to manage ports 2379 and 2380. dqlite communication happens over port 9345.
- Unified Tooling: Backups are handled via the K3s binary itself rather than a separate
etcdctlinstallation.
Practical Implementation: Deploying an HA Cluster
To deploy a dqlite-backed HA cluster, you initialize the first node as the cluster seed and join subsequent nodes to it. This assumes you are using K3s v1.19 or later on Linux nodes with root or sudo permissions.
Step 1: Initialize the first server
Run this on your first node to start the embedded datastore:
curl -sfL https://get.k3s.io | sh -s - server --cluster-init
Step 2: Join additional servers
On the second and third nodes, join the cluster using the first node's IP. Replace <FIRST_NODE_IP> with the actual IP address:
curl -sfL https://get.k3s.io | sh -s - server --server https://<FIRST_NODE_IP>:6443
Step 3: Verify HA Status
Run the following command on any server node to confirm all nodes are acting as the control plane:
k3s kubectl get nodes -o wide
Expected result: All server nodes should show the control-plane role and a Ready status.
Testing Failover and Recovery
The primary reason to use dqlite is to survive a node failure. You can verify the Raft election process by simulating a crash on the current leader.
- Identify the leader by checking the
kube-controller-managerendpoints in thekube-systemnamespace. - Stop the K3s service on the leader node:
systemctl stop k3s. - Watch the cluster events on a remaining node:
k3s kubectl get events -A --watch.
You should observe a new leader being elected within 5 to 10 seconds, and the API server should remain responsive.
Trade-offs: When to Avoid dqlite
dqlite is optimized for footprint and simplicity, not maximum throughput. It is a poor choice for environments with extreme "churn"—where thousands of objects (like ConfigMaps or Secrets) are created and deleted every second by GitOps controllers or heavy automation.
Key Limitations:
- Write Latency: Under heavy load, dqlite may exhibit higher latency than a tuned, external etcd cluster.
- No Direct Inspection: You cannot use
etcdctlto query the database; you must rely onkubectlor K3s snapshots. - Scaling Ceiling: For clusters exceeding 100 nodes or those with massive CRD usage, an external PostgreSQL or etcd database is recommended.
Operational Safety: Backups and Rollbacks
Because dqlite is embedded, you cannot simply copy the database file between nodes (as Raft membership is encoded in the data). Instead, use the built-in snapshot utility.
To save a snapshot:
k3s etcd-snapshot save --name pre-upgrade-backup
Snapshots are stored in /var/lib/rancher/k3s/server/db/snapshots/.
To rollback/restore:
If the datastore is corrupted or a migration fails, stop all K3s servers and run the reset command on one node:
k3s server --cluster-reset --cluster-reset-restore-path=/path/to/snapshot
After the reset, the other nodes can re-join the cluster using the --server flag.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.