Finding the Optimal Concurrency Value
The concurrency value that yields the lowest 95th-percentile (p95) latency is the highest possible value that does not trigger CPU saturation or memory exhaustion for your specific workload. Because there is no universal optimum, you must identify the "knee" of the latency curve where the benefit of reduced queueing (from higher concurrency) is offset by the cost of CPU contention (from too many simultaneous requests).
The Impact of Doubling CPU Allocation
When CPU allocation is doubled, the optimal concurrency setting generally increases, but not necessarily linearly. Doubling CPU provides more compute cycles to share among concurrent requests, which shifts the CPU-throttling threshold higher. This allows you to increase concurrency to further reduce queueing delays and cold starts without increasing the per-request processing time.
Distinguishing Queueing from Throttling
To determine why latency is spiking during bursts, monitor the following signals in Cloud Monitoring:
- Queueing Latency: Indicated by a spike in
container/instance_count lagging behind the request rate, combined with high request_latency but low container/cpu/utilization. This suggests requests are waiting for a free slot or a new instance to boot.
- CPU Throttling: Indicated by
container/cpu/utilization hitting 100% (or the allocated limit) while container/instance_count is stable. In this state, the CPU is saturated, and requests are competing for cycles, increasing the processing time for every active request.
Verification Steps
To validate your settings, perform a scoped load test using a tool like k6 or hey:
- Deploy the service with a low concurrency setting (e.g., 10) and send a burst of requests exceeding that limit. Record the p95 latency.
- Increase concurrency (e.g., to 80) and repeat the burst.
- Compare the
container/cpu/utilization metric during the burst. If CPU utilization is low but latency is high, increase concurrency. If CPU utilization is maxed out and latency rises, decrease concurrency or increase CPU allocation.
Missing Diagnostic: To refine this recommendation, what is the primary nature of your workload (I/O-bound vs. CPU-bound)? I/O-bound services can typically sustain much higher concurrency than CPU-bound services before hitting the throttling threshold.