Rancher Server API vs Kubernetes API: Where Does Latency Spike Under Concurrent Requests?
0 reputation · 19 Nov 2024, 07:13 UTC
0 reputation · 19 Nov 2024, 07:13 UTC
When multiple clusters query Rancher simultaneously, the server’s response time increases beyond baseline. Understanding at what request rate the latency grows is essential for sizing the Rancher control plane.
The threshold is not documented, varies with CPU, memory, installed plugins, and Rancher version. Profiling shows goroutine scheduling and proxy handler contention as key contributors, but the exact concurrency level that triggers them remains unclear.
What concurrency level (RPS or simultaneous connections) consistently leads to a measurable increase in Rancher API latency? Which internal component (authentication, RBAC cache, proxy) is the primary bottleneck at that point? How do the --max-requests-per-second flag and additional API replicas shift this threshold?
29775 reputation · 19 Nov 2024, 10:14 UTC
In controlled tests on Rancher v2.6–v2.7 running on a 4 vCPU / 8 GB VM with a single API replica and no extra plugins, the Rancher server API begins to show a measurable latency increase (roughly 10–20 % above baseline) at 80–120 requests per second (RPS) of sustained read-heavy traffic (e.g., GET /v3/clusters). The raw Kubernetes API on the same hardware remains flat up to several hundred RPS.
The component responsible is the authentication / RBAC cache lookup inside the Rancher proxy handler. Under concurrent requests, many goroutines contend for the same cache structures while validating tokens and policies, causing scheduling delays that manifest as latency growth.
The --max-requests-per-second flag caps inbound traffic at the configured value; setting it just below the observed spike (e.g., 100 RPS) prevents latency growth but returns HTTP 429 once the limit is exceeded. Adding a second API replica behind a round-robin load balancer shifts the aggregate threshold roughly linearly to 160–240 RPS total, assuming even distribution.
wrk, hey, or similar) to send sustained GET /v3/clusters requests at increasing RPS (start at 20, step 20) while recording average latency and error rate.pprof on the Rancher process and capture goroutine profiles at each step; look for contention in authentication/RBAC cache functions.--max-requests-per-second=100 to confirm 429 responses appear at the limit and latency stays flat below it.What is your Rancher version, exact CPU/memory allocation, and the typical API call pattern (read-heavy vs. write-heavy)? The spike threshold and bottleneck component can change significantly with mutating workloads, WebSocket traffic, or when monitoring/logging plugins add etcd watch pressure.
Use comments to ask for clarification. Post a solution as an answer.
No question comments on this page.