Diagnosing k3OS Automatic OS Upgrade Failures on Bare‑Metal Clusters
Step‑by‑step guide to diagnose why k3OS automatic OS upgrades fail on bare‑metal nodes, covering service status, logs, network, disk space, cloud‑config, and hardware compatibility.
31 Aug 2025, 03:41 UTC

Recognizable Condition
After a scheduled k3OS upgrade, nodes fail to reboot into the new OS version, remain on the old release, or show a kernel panic during boot.
Cause/Diagnostic Table
| Symptom | Likely Cause |
|---|---|
| k3os-upgrade service inactive or failed | Service did not start; upgrade process never began |
| Download timeouts in logs | Network connectivity to upgrades.rancher.com blocked |
| "No space left on device" messages | Insufficient free space on /var/lib/k3os for image extraction |
| YAML syntax error in cloud‑config | Malformed user‑data causing upgrade script abort |
| Kernel panic after reboot | Incompatible drivers or missing firmware for hardware |
Ordered Checks
- Verify the upgrade service state
Run
systemctl status k3os-upgradeas root. Look foractive (exited)orfailed. If the unit isinactiveor shows a failure, note the error message. - Inspect the upgrade log
Examine
/var/log/k3os-upgrade.logfor tracebacks, HTTP errors, or space‑related messages. Usejournalctl -u k3os-upgrade --no-pageras an alternative. - Test outbound HTTPS to the upgrade server
From the node, run
curl -v https://upgrades.rancher.com/v1. A successful response returns HTTP 200 and a JSON payload. Note any connection timeout, TLS handshake failure, or HTTP error codes. - Check free space on the upgrade partition
Execute
df -h /var/lib/k3os. Ensure at least 2 GB of free space is reported. If the usage is >80 % or free space <2 GB, the upgrade cannot extract the new image. - Validate the cloud‑config user‑data
If the node uses a custom cloud‑config, retrieve the current user‑data (often stored in /var/lib/rancher/k3s/server/manifests or provided via DHCP). Run a YAML linter such as
yamllintor an online validator. Look for indentation errors, forbidden keys, or missing required fields. - Review hardware compatibility
If logs indicate a kernel panic, compare the kernel version in the new release with the hardware’s driver/firmware requirements. Check dmesg output after a failed boot for messages like
Unable to handle kernel paging requestorrequest_firmware failed.
Fixes Tied to Findings
- Service not starting – Enable and start the unit:
systemctl enable --now k3os-upgrade. If the unit file is missing, reinstall the k3os package (apk add --upgrade k3os) on the node. - Network blockage – Ensure outbound TCP 443 to upgrades.rancher.com is allowed. Adjust firewall rules or proxy settings. Verify DNS resolution with
nslookup upgrades.rancher.com. - Insufficient disk space – Clean the upgrade partition: remove old cached images (
rm -rf /var/lib/k3os/*) or expand the partition if using LVM. After freeing space, restart the upgrade service. - Cloud‑config syntax error – Correct the YAML, then reapply the configuration (e.g., via
ros config setor by updating the user‑data source and triggering a reconfigure). After correction, trigger the upgrade again withsystemctl restart k3os-upgrade. - Kernel panic / driver mismatch – Boot into the previous working kernel (often available via the GRUB entry), then either:
- Upgrade firmware/drivers on the hardware, or
- Pin the k3OS version to a known‑good release by setting
K3OS_UPGRADE_ENABLED=falsein /etc/rancher/k3os/config.yaml until compatibility is verified.
Escalation Criteria
Proceed to escalation when:
- The upgrade service repeatedly fails after applying the above fixes and logs show persistent "internal server error" or "502 Bad Gateway" from the upgrade endpoint.
- Disk space cannot be freed because the partition is at its maximum size and resizing is not possible without reinstall.
- Kernel panic persists after firmware updates and the hardware is not listed in the k3OS hardware compatibility matrix for the target version.
- Multiple nodes in the cluster exhibit the same failure pattern, suggesting a systemic issue with the upgrade server or network path.
In these cases, collect the following data before contacting Rancher support or the community:
- Output of
systemctl status k3os-upgradeandjournalctl -u k3os-upgrade - Contents of
/var/log/k3os-upgrade.log - Result of
curl -v https://upgrades.rancher.com/v1 - Output of
df -h /var/lib/k3osandlsblk - Current cloud‑config/user‑data YAML
- Hardware model, BIOS/firmware versions, and output of
dmesgfrom a failed boot.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.