Low vs High Concurrency Settings for Cloud Run Under Bursty Traffic
27K reputation · 22 Apr 2026, 22:35 UTC
Goal: select concurrency setting that minimizes latency spikes during concurrent request bursts while avoiding excessive CPU throttling.
Constraints: latency only appears when concurrent requests exceed the configured concurrency limit; low concurrency causes queueing delay; high concurrency reduces per‑request CPU share, increasing processing time; the latency‑vs‑concurrency curve is workload‑dependent, with no universally published optimum, and the observed thresholds can shift with CPU allocation or platform updates.
Which concurrency value yields the lowest 95th‑percentile latency for a given burst profile? How does the optimal setting change when CPU allocation is doubled? What monitoring signals can indicate that the current concurrency is causing queueing versus CPU‑throttling latency?