State Synchronization Mechanism
Rancher does not utilize a traditional linear "reconciliation queue" for pending changes during an agent outage. Instead, it employs a declarative state model. The Rancher Manager maintains the desired state of the managed cluster in its own database, while the Cluster Agent is responsible for ensuring the downstream cluster matches that desired state.
Reconnection and Prioritization
When a Cluster Agent reconnects after a prolonged outage, it does not replay a sequence of historical events. Instead, it performs a full synchronization of the current desired state. The prioritization follows these general principles:
- Latest State Wins: The agent fetches the most recent version of the configuration from the Manager. Any intermediate changes made while the agent was offline are effectively collapsed into the final desired state.
- Resource Dependency: Reconciliation typically follows Kubernetes' native controller patterns, where foundational resources (Namespaces, CRDs) are processed before dependent workloads.
- Periodic Polling: The agent continuously polls the Manager to ensure the local state converges with the remote state, correcting any drift caused by the disconnection.
Likely Causes of Inconsistency
While the declarative model prevents "out-of-order" execution of old commands, inconsistencies during Kubernetes upgrades usually stem from:
- Version Mismatch: If the Kubernetes API version changed during the outage, the agent may struggle to apply manifests that are no longer compatible with the upgraded cluster.
- Manual Intervention: Changes made directly to the downstream cluster via
kubectl during the outage may conflict with the Manager's desired state, leading to "flapping" resources as the agent attempts to revert them.
Verification Steps
To verify the synchronization status after reconnection, check the agent logs and the cluster condition in the Rancher UI. You can manually inspect the agent's connectivity and state via the following scoped commands:
# Check the status of the cluster-agent pods in the cattle-system namespace
kubectl get pods -n cattle-system -l app=cattle-cluster-agent
# View logs for synchronization errors or API incompatibilities
kubectl logs -n cattle-system -l app=cattle-cluster-agent --tail=100
Diagnostic Request: To provide a more specific resolution, please clarify if the inconsistency is appearing as a Waiting state in the Rancher UI or as CrashLoopBackOff events within the cattle-system namespace.