Diagnosing k3OS Boot Stalls Due to Missing or Malformed cloud‑init User‑Data
Learn how to identify and fix k3OS nodes that halt at the initramfs prompt because cloud‑init cannot locate or parse its user‑data.
26 Jul 2026, 21:53 UTC

Recognizable condition
The k3OS node stops before the k3s service starts. You see either an initramfs prompt or a repeating message such as "Waiting for cloud‑init data source" on the console or serial output. The system never proceeds to load the k3s container runtime.
Cause / diagnostic table
| Possible cause | What to look for |
|---|---|
| Missing cloud‑init configuration file | /var/lib/rancher/k3os/cloud‑init.yml does not exist |
| Malformed YAML in cloud‑init | File exists but contains syntax errors (e.g., incorrect indentation, invalid anchors) |
| Network not ready when cloud‑init runs | DHCP lease not obtained; cloud‑init logs show timeout waiting for network |
Ordered checks
- Examine boot logs
On the node console or serial connection, look for cloud‑init messages. Typical lines start with "cloud-init" and include "Cloud‑init v." followed by either "finished" or an error.
# Run on the node (requires root or console access) journalctl -u cloud-init | grep -i cloud-init - Check for the cloud‑init file
Verify whether the expected user‑data file is present.
# On the node (root) ls -l /var/lib/rancher/k3os/cloud-init.ymlIf the file is missing, note the absence; if present, proceed to validate its content.
- Validate YAML syntax
Use a YAML parser to ensure the file is well‑formed.
# On the node (root) – python3 is usually available python3 -m yaml.tool /var/lib/rancher/k3os/cloud-init.yml 2>&1 # Alternatively, if yamllint is installed: yamllint /var/lib/rancher/k3os/cloud-init.ymlAny output indicating an error points to a syntax problem.
- Test network connectivity
Confirm that the node can obtain an IP address via DHCP before cloud‑init runs.
# On the node (root) dhclient -v eth0 # replace eth0 with the appropriate interface if needed # Or, if a static address is expected, verify it is configured: ip addr show eth0Look for a successful "bound to" message or a configured address.
Fixes tied to findings
- Missing cloud‑init file
- If you have a backup of the intended user‑data, copy it back:
# On the node (root) cp /path/to/backup/cloud-init.yml /var/lib/rancher/k3os/cloud-init.yml chmod 644 /var/lib/rancher/k3os/cloud-init.yml- Otherwise, re‑apply the user‑data via the k3os config URL (usually set in
/etc/k3os/config.yaml) and reboot:
# Edit the config to point to the correct source (requires reboot) vi /etc/k3os/config.yaml # Ensure the 'cloud_init_url' field is correct reboot - Invalid YAML
- Correct the syntax errors reported by the validator, then save the file.
# Example fix: correct indentation vi /var/lib/rancher/k3os/cloud-init.yml # After editing, validate again: python3 -m yaml.tool /var/lib/rancher/k3os/cloud-init.yml - Network not ready
- Ensure a DHCP server is reachable on the network segment.
- If DHCP cannot be used, configure a static network in
/etc/k3os/config.yamlunder thenetworksection.
# Example static config (edit then reboot) vi /etc/k3os/config.yaml network: interfaces: eth0: address: 192.168.1.42/24 gateway: 192.168.1.1 dns: - 8.8.8.8 reboot
Escalation criteria
- After applying the appropriate fix, reboot the node.
- If the node still does not reach the k3s service after two consecutive reboots, collect the following logs for further analysis:
# On the node (root)
cp /var/log/cloud-init.log /tmp/cloud-init.log
cp /var/log/messages /tmp/messages.log
# Transfer the logs off the node (e.g., via scp) for inspection
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.