Tuning Cloud Run Concurrency and Scaling to Balance Cost and Performance
Learn how to tune Cloud Run concurrency and scaling limits to prevent OOM crashes and control costs while eliminating cold starts.
02 Sept 2026, 14:25 UTC

The Trade-off Between Scaling and Resource Exhaustion
When deploying to Cloud Run, the default scaling behavior often leads to one of two problems: either too many container instances spin up, driving up costs, or a single instance attempts to handle more requests than its memory can support, leading to Out-of-Memory (OOM) crashes.
The key to solving this is the relationship between Concurrency (how many requests one instance handles at once) and Max Instances (the ceiling for horizontal scaling). Tuning these prevents "runaway scaling" during traffic spikes while ensuring your application remains responsive.
Prerequisites
- A Google Cloud Project with the Cloud Run API enabled.
- A containerized application deployed to Artifact Registry.
gcloudCLI installed and authenticated withroles/run.adminpermissions.
Configuring Concurrency and Scaling Limits
Concurrency defines the threshold at which Cloud Run starts a new instance. If concurrency is set to 80 and 81 requests arrive simultaneously, Cloud Run spins up a second instance to handle the 81st request.
Run the following command in your terminal to update an existing service. Replace [SERVICE_NAME] and [REGION] with your specific deployment details:
gcloud run services update [SERVICE_NAME] \
--region [REGION] \
--concurrency 50 \
--max-instances 10 \
--min-instances 1
Configuration Breakdown
--concurrency 50: Each instance will handle up to 50 simultaneous requests. This is ideal for I/O-bound applications (like Node.js or Go). For CPU-heavy tasks, lower this number to avoid latency spikes.--max-instances 10: Limits the total number of containers to 10. This acts as a cost circuit-breaker, preventing an unexpected traffic surge from exhausting your budget.--min-instances 1: Keeps one instance "warm" at all times. This eliminates the cold start (the delay while a new container boots) for the first request.
CPU Allocation Strategies
Cloud Run offers two CPU allocation modes that fundamentally change how your app behaves between requests.
| Setting | Behavior | Best Use Case |
|---|---|---|
| CPU allocated during request processing | CPU is throttled to near zero immediately after the response is sent. | Standard REST APIs, lightweight webhooks. |
| CPU always allocated | CPU remains available even when no requests are active. | Background processing, asynchronous tasks, WebSocket connections. |
To switch to always-allocated CPU, use the flag --no-cpu-throttling during deployment. Risk: This increases costs because you are billed for the instance's entire lifetime, not just the request duration.
Verification and Diagnostics
To verify your scaling configuration, you can simulate a load spike using a tool like hey or ab (Apache Benchmark). If you set concurrency to 1, every single simultaneous request should trigger a new instance.
Diagnostic Check:
- Deploy with
--concurrency 1. - Send 5 simultaneous requests to the service URL.
- Navigate to the Cloud Run Console > [Service] > Revisions tab.
- Observe the "Instance Count" metric; it should spike to 5.
Rollback and Recovery
If you set concurrency too high and notice 504 Gateway Timeout errors or OOMKilled events in the logs, you must reduce the concurrency limit immediately to force the load to spread across more instances.
# Recovery: Reduce concurrency to spread load and increase stability
gcloud run services update [SERVICE_NAME] --region [REGION] --concurrency 10
Limitations
- Memory Pressure: High concurrency does not magically increase memory. If each request consumes 100MB of RAM and you set concurrency to 100 on a 2GB instance, the container will crash.
- Cold Start Latency: While
min-instancessolves cold starts for the baseline, any traffic exceedingmin-instances * concurrencywill still trigger cold starts for new instances.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.