1. Quick answer
k3s does not expose a direct flag to increase the watch‑history TTL. The 410 Gone you see is caused by the continue token becoming stale (usually after a server restart or a long pause between pages). The practical fix is to keep the limit moderate, retry immediately on 410, and, if possible, avoid long pauses or frequent restarts.
2. Confirmed facts
- The k8s API server (used by k3s) returns 410 Gone when it cannot resolve a
continue token.
- Tokens expire after a short period or after a server restart; the default TTL is not configurable via a k3s flag.
- Large
limit values (>500) can increase memory pressure and accelerate token expiry.
3. Likely explanation
- During rapid dataset growth, the first page’s
resourceVersion can become stale before the next request is issued.
- A stale
continue token triggers a 410 Gone, indicating the historical state is no longer available.
- This is a normal API‑server behaviour, not a bug in k3s.
4. Practical steps for this case
- Use a moderate limit – 200 to 500 items per page. Avoid
limit=1000 or higher.
- Retry immediately on 410 – if you receive a 410, drop the
continue token and issue a fresh GET with the same limit. The API server will start from the current state.
- Throttle pauses – keep the interval between paginated requests short (seconds, not minutes). Long pauses give the token time to expire.
- Monitor restarts – if the API server restarts frequently, consider reducing the restart frequency or adding a small back‑off between pages.
- Check logs for token expiry – look for “continue token expired” or “cannot resolve token” in
journalctl -u k3s or /var/log/k3s.log to confirm the cause.
5. No watch‑list flag in k3s
k3s does not provide an annotation or flag that lets a client request the initial List as part of a Watch call. The standard Kubernetes API requires a separate GET for the initial list, followed by a watch that streams changes from that resourceVersion. This design is intentional to keep the watch API lightweight.
6. Best‑practice recommendations
- Use the client‑go
ListOptions{Limit: 200} and Continue pattern; client‑go handles 410 automatically by retrying the list.
- When writing custom code, catch 410, clear the
continue token, and re‑issue the list.
- For very large datasets, consider server‑side filtering (
labelSelector, fieldSelector) to reduce the number of items returned.
- Keep the cluster’s
k3s binary up to date; newer releases may adjust token TTL or improve log verbosity.
7. Diagnostic detail needed
To tailor the recommendation, could you confirm whether the API server restarts frequently (e.g., due to health‑check failures or rolling upgrades)? If so, a brief pause between paginated requests might be necessary.