Rancher Server API vs Kubernetes API: Where Does Latency Spike Under Concurrent Requests?
0 reputation · 19 Nov 2024, 07:13 UTC
When multiple clusters query Rancher simultaneously, the server’s response time increases beyond baseline. Understanding at what request rate the latency grows is essential for sizing the Rancher control plane.
The threshold is not documented, varies with CPU, memory, installed plugins, and Rancher version. Profiling shows goroutine scheduling and proxy handler contention as key contributors, but the exact concurrency level that triggers them remains unclear.
What concurrency level (RPS or simultaneous connections) consistently leads to a measurable increase in Rancher API latency? Which internal component (authentication, RBAC cache, proxy) is the primary bottleneck at that point? How do the --max-requests-per-second flag and additional API replicas shift this threshold?