Diagnosing Talos Linux Nodes Stuck in NotReady: API Connectivity and Config Mismatches
A step‑by‑step diagnostic guide for Talos Linux nodes stuck in NotReady, covering endpoint misconfiguration, TLS errors, firewall blocks, resource pressure, and network‑interface mismatches with concrete commands and verification steps.
22 May 2026, 03:16 UTC

Recognizable condition
Running kubectl get nodes shows one or more Talos nodes with a status of NotReady. The node may have joined the cluster but the kubelet cannot maintain a healthy connection to the Kubernetes API server or the Talos API.
Cause / diagnostic quick‑reference table
| Observed symptom | Likely cause | Quick verification |
|---|---|---|
Node reports NotReady immediately after boot |
Cluster endpoint (API server address) wrong in machine config | talosctl get machineconfig | grep -A2 endpoint |
Kubelet logs show x509: certificate signed by unknown authority |
TLS certificate mismatch or expiration between node and control plane | talosctl logs kubelet | grep -i cert |
| Connection timeouts to port 6443 or 50000 | Firewall / load‑balancer blocking Kubernetes API (6443) or Talos API (50000) | nc -zv 6443 and nc -zv 50000 |
Node shows DiskPressure or MemoryPressure conditions |
Insufficient host resources causing kubelet/containerd crashes | talosctl stats or df -h /var/lib/kubelet on the node console |
| Machine config applies but node reboots repeatedly | Incorrect network interface name or duplicate MAC in network.interfaces |
talosctl get machineconfig | grep -A5 interfaces |
Ordered diagnostic checks
-
Validate the cluster endpoint in the machine configuration
Run from your admin workstation (where
talosctlis configured with the Talos client certificate):talosctl get machineconfig --output yaml | yq '.cluster.endpoint'Expected: a reachable IP or DNS name that resolves to the control‑plane load balancer. If the value is
https://127.0.0.1:6443or an outdated address, the kubelet cannot authenticate.Fix: Edit the machine config (or the cluster‑level
cluster.yaml) and set the correct endpoint, then apply:talosctl apply-config --nodes --file corrected-machineconfig.yamlRisk:
apply-configtriggers a full node reboot. Schedule during a maintenance window. -
Inspect kubelet logs for authentication or networking errors
talosctl logs kubelet --followLook for lines containing
Failed to get node info,x509, ordial tcp. Those point directly to TLS or connectivity problems.Fix: If TLS errors appear, regenerate the node’s client certificate via
talosctl gen cert --nodeand re‑apply the config. -
Verify network reachability to the API server and Talos API ports
From the same workstation (or a bastion host in the same VPC):
nc -zv 6443 nc -zv 50000Both should succeed. If they fail, check security groups, firewall rules, and the load‑balancer health checks.
Fix: Add inbound rules for TCP 6443 (Kubernetes API) and TCP 50000 (Talos API) to the control‑plane security group.
-
Check TLS certificate validity on the node
On the node console (accessible via
talosctl consoleor serial console):openssl x509 -in /var/lib/kubelet/pki/kubelet-client-current.pem -text -noout | grep -E 'Not Before|Not After|Subject:'If the
Not Afterdate is past, the certificate has expired.Fix: Rotate certificates with
talosctl rotate certs --nodes(requires control‑plane quorum). -
Assess host resource pressure
talosctl stats # or from node console df -h /var/lib/kubelet /var/lib/containerd free -mDisk usage > 85% or memory < 10% free often leads to kubelet OOM kills.
Fix: Clean up unused container images (
crictl rmi --prune), expand the root filesystem, or add a larger disk. -
Confirm machine‑config network interface mapping
talosctl get machineconfig --output yaml | yq '.network.interfaces[] | select(.device=="eth0")'Ensure the
devicename matches the actual NIC (e.g.,ens5on AWS Nitro). A mismatch prevents the node from acquiring an IP on the cluster network.Fix: Update the interface definition and re‑apply the config (see step 1).
Fixes tied to findings
| Finding | Remediation | Verification after fix |
|---|---|---|
Wrong cluster.endpoint |
Edit machine config, run talosctl apply-config |
kubectl get nodes -o jsonpath='{.status.conditions[?(@.type=="Ready")].status}' returns True |
| TLS certificate error | talosctl rotate certs --nodes |
Kubelet logs no longer show x509 errors; node becomes Ready |
| Port 6443/50000 blocked | Update security group / firewall rules | nc -zv 6443 succeeds; talosctl health shows OK |
| Disk / memory pressure | Prune images, expand volume, or add swap | talosctl stats shows diskPressure: false and memoryPressure: false |
| Interface name mismatch | Correct network.interfaces[].device and re‑apply |
Node obtains expected IP; ip addr show on console matches config |
Escalation criteria
- Control‑plane quorum lost (fewer than
(n/2)+1healthy control‑plane nodes). In this state,talosctl rotate certsorapply-configon control‑plane nodes will fail. - Repeated node reboots after
apply-configwithout reaching Ready – indicates a deeper hardware or firmware issue. - Certificate rotation fails because the CA private key is unavailable – requires cluster‑level disaster recovery.
When any of the above conditions are met, engage the platform operations team or open a support ticket with the Talos maintainers. Provide the output of talosctl health, talosctl logs kubelet , and the current machine config for rapid triage.
Limitations & practical verification
This guide covers the most common API‑connectivity and configuration mismatches that drive a Talos node into NotReady. It does not address:
- Underlying hypervisor / cloud‑provider networking bugs (e.g., ENI attachment failures).
- Custom CNI misconfiguration that manifests after the node is Ready.
- Hardware faults such as NIC failure or disk corruption.
Final verification step: After applying the relevant fix, run:
kubectl get nodes -o wide
The STATUS column should read Ready and the CONDITIONS should show Ready=True with no DiskPressure or MemoryPressure. If the node remains NotReady after all checks, treat it as an escalation case.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.