Answer first
For a low‑resource edge node that values minimal configuration and can tolerate brief downtime, run k3s with its default embedded SQLite store, add a periodic integrity‑check and a backup/restore workflow. Switch to an external etcd cluster only when the node needs high availability or sustained write loads that exceed SQLite’s practical limits.
Confirmed facts
- k3s stores all API objects in a single SQLite file (
k3s.db) by default. If that file becomes corrupted the k3s server process exits and there is no built‑in repair.
- SQLite provides a
PRAGMA integrity_check; command that returns "ok" for a healthy file and an error message otherwise.
- Restoring a known‑good copy of
k3s.db (from a file backup or volume snapshot) and restarting k3s is the only supported recovery path.
- External etcd enables HA and higher write throughput but requires operating a separate etcd cluster (typically 3‑5 nodes) and additional networking configuration.
Steps for SQLite‑based edge node
- Enable integrity monitoring – create a systemd timer (or cron job) that runs:
sqlite3 /var/lib/rancher/k3s/server/k3s.db "PRAGMA integrity_check;"
Treat any non‑"ok" output as an alert (e.g., log to journald and trigger a notification).
- Backup strategy – while k3s is stopped (or using a quiescent volume snapshot), copy
k3s.db to a secure location. Frequency should match the node’s change rate; a common starting point is every 15‑30 minutes for modest workloads.
- Recovery procedure – upon an integrity‑check failure:
- Stop k3s:
systemctl stop k3s
- Replace the corrupted file:
cp /path/to/backup/k3s.db /var/lib/rancher/k3s/server/k3s.db
- Start k3s:
systemctl start k3s
This restores the API state to the point of the last good backup; no manual recreation of the database is needed.
When to consider external etcd
Switch to an external etcd cluster if any of the following apply:
- The node must sustain write loads above roughly 500 operations per second (where SQLite file‑lock contention becomes noticeable).
- Downtime of more than a few seconds is unacceptable (e.g., telemetry or control loops that cannot tolerate API server restarts).
- Operational overhead of managing a small etcd cluster is justified by the HA requirement.
In that case, provision a 3‑node etcd cluster, point k3s to it with --datastore-endpoint, and disable the embedded SQLite store.
Missing diagnostic detail
To finalize the recommendation, please confirm the expected average write load (operations per second) or the maximum tolerable downtime for your edge workload. This information determines whether the SQLite‑based approach remains suitable or if etcd is required.