Rancher Cluster Agent Connectivity: Diagnosing WebSocket Failures
When Rancher agents show CrashLoopBackOff with WebSocket handshake errors, follow this diagnostic guide to identify network, TLS, proxy, or token issues and apply targeted fixes.
20 Sept 2026, 13:30 UTC

The Problem: Agents Stuck in CrashLoopBackOff
When downstream clusters appear Unavailable or Waiting in the Rancher UI, the cattle-cluster-agent and cattle-node-agent pods are often in CrashLoopBackOff. The root cause is typically a WebSocket connection failure to wss://<rancher-url>/v3/connect. This guide helps you diagnose and fix the issue systematically. Commands assume you have cluster-admin access to the downstream cluster via kubectl and network access from a cluster node.
Quick Diagnostic Table
| Symptom | Most Likely Cause |
|---|---|
| CrashLoopBackOff + bad handshake | Proxy/LB missing WebSocket headers or low timeout |
| CrashLoopBackOff + x509 error | TLS certificate mismatch or missing CA |
| Error=connection refused | Network policy or firewall blocking port 443 |
| Waiting > 15 min, no logs | Registration token expired or agent version skew |
Step-by-Step Checks
- Verify Rancher server reachability
Run from a node in the downstream cluster (replace<rancher-url>with your server hostname):curl -vk https://<rancher-url>/v3/connect
Expected: an HTTP response indicating a WebSocket upgrade endpoint. If the connection times out or is refused, check network connectivity first. - Inspect agent logs
kubectl logs -n cattle-system -l app=cattle-cluster-agent --tail=100
Look forfailed to connect to wss://,websocket: bad handshake, orx509: certificate signed by unknown authority. - Check pod events
kubectl describe pod -n cattle-system -l app=cattle-cluster-agent
Events likeFailedMountor Pod Security admission denials point to RBAC or PSA issues. - Validate TLS trust
openssl s_client -connect <rancher-url>:443 -servername <rancher-url>
Verify the certificate chain is trusted and the hostname matches the SAN list. - Confirm registration token validity
kubectl get secret -n cattle-system cattle-credentials -o jsonpath='{.data.value}' | base64 -dAn empty or malformed token indicates the cluster needs re-registration. - Review network policies
kubectl get networkpolicy -A -o wide
Ensure no policy blocks egress to port 443 on the Rancher server IP. - Check proxy/load balancer config
For NGINX or HAProxy in front of Rancher, verify:proxy_set_header Upgrade $http_upgrade;proxy_read_timeout 300s;- The
Connection: upgradeheader is passed through
Applying Fixes
- Network/Firewall: Open outbound TCP 443 to the Rancher server from all cluster nodes.
- TLS Certificate Issues: Add the Rancher CA certificate to the cluster's trusted store, or reinstall the agent with the CA provided (for example
--set rancherServerCA=<ca-content>during registration). - Proxy/LB Misconfiguration: Enable WebSocket upgrade headers and increase the idle timeout to at least 60 seconds (300s recommended).
- Token Expired: Delete the
cattle-credentialssecret and re-register the cluster via the Rancher UI. Schedule this during a maintenance window; re-registration can briefly disrupt Rancher-managed resources. - Version Skew: Align the agent image tag with the Rancher server version. Changing the image on an existing cluster requires a helm upgrade of the agent release, not just a pod restart.
- RBAC/PSA Blocking: Grant the cattle-system namespace the required Pod Security Admission level (privileged) so the node-agent can use hostPath mounts.
When to Escalate to SUSE
Escalate if:
- Network and TLS checks pass but
bad handshakepersists after 30 minutes. - The cluster remains in Waiting state with no agent logs.
- Failure occurs immediately after a Rancher server upgrade.
- Custom CNI plugins appear to block WebSocket upgrade traffic.
- The environment has FIPS or air-gapped restrictions preventing standard diagnostics.
Verification Steps
- Test agent connectivity from inside the pod
kubectl exec -n cattle-system -it deploy/cattle-cluster-agent -- wget -qO- https://<rancher-url>/v3/connect
Should return a WebSocket upgrade response rather than a connection error. - Check cluster state via the Rancher API
curl -k -H 'Authorization: Bearer <token>' https://<rancher-url>/v3/clusters/<cluster-id>
Verifystate: "active". - Validate version alignment
kubectl get pods -n cattle-system -o jsonpath='{.items[*].spec.containers[*].image}'Ensure the agent image minor version matches the Rancher server. - Monitor for certificate errors
kubectl logs -n cattle-system -l app=cattle-cluster-agent --since=30m | grep -iE 'cert|tls|x509'
Should return no results over a 30-minute window.
Caution: Do not delete the cattle-system namespace on production clusters without a backup; it removes Rancher-managed resources. Test NetworkPolicy changes in staging first, since they can affect other workloads.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.