Cloud Run Concurrency Tuning: A Practical Guide to Cost and Performance
Concurrency is Cloud Run's biggest lever for cost and performance. Learn how to calculate the right value for your workload, avoid OOM kills, and verify your tuning with actual metrics.
31 Jul 2025, 00:54 UTC

The Core Problem: Cost vs. Cold Start Trade‑off
Concurrency is Cloud Run’s single most impactful setting for balancing cost, performance, and reliability. The default of 80 concurrent requests per instance works for many services, but picking the wrong value can cause OOM kills, excessive cold starts, or unexpected bills.
How Concurrency Works
Each Cloud Run instance handles up to containerConcurrency simultaneous requests before the autoscaler spins up another instance. The autoscaler aims for roughly 60% utilization, so the expected instance count is approximately
ceil(RPS × avgLatency / (concurrency × 0.6))
Worked Example: Node.js API
Assume a Node.js API with the following characteristics:
- Average request latency: 200 ms
- Memory per request: ~150 MB
- Target throughput: 40 RPS
Setting concurrency to 40 means one instance can handle about 12 RPS (40 × 60% ÷ 0.2 s). To serve 40 RPS you need
ceil(40 ÷ 12) = 4 instances
Each instance would use roughly 6 GB of memory (40 × 150 MB). A 1 GB instance limit would fail, so raise the memory to 8 GB or lower concurrency.
Configuration Example
# Set concurrency to 40 and memory to 8 Gi for a service named 'api-service'
gcloud run services update api-service --concurrency 40 --region us-central1 --memory 8Gi
When to Adjust Concurrency
Decrease concurrency (1–10) when:
- Using single‑threaded runtimes (Python sync, Ruby)
- Each request consumes significant memory (JVM)
- Handling long‑lived connections (WebSockets, streaming)
- Observing frequent OOM kills in logs
Increase concurrency (100–500) when:
- Requests are lightweight and fast (<50 ms)
- Memory per request is low (<50 MB)
- Want to minimize instance count and cold starts
- Using async‑capable runtimes (Node.js with async/await)
Common Mistakes and How to Avoid Them
- Mismatched memory and concurrency: Setting concurrency to 100 with 256 MB per request and a 1 GB limit guarantees OOM kills. Rule of thumb: concurrency × memoryPerRequest ≤ memoryLimit × 0.8.
- Ignoring runtime characteristics: A Python Flask app with 80 concurrent requests will thrash the GIL. Use concurrency 1–10 or switch to async frameworks.
- Not testing at scale: Single‑client tests hide queuing and instance scaling. Use distributed load tools.
- Long‑lived connections: WebSockets or streaming endpoints occupy a slot for minutes. Reserve low concurrency for such services.
Verification Checklist
- Run a 5‑minute load test at expected peak RPS using
heyork6. - Monitor
run.googleapis.com/container/instance_countin Cloud Monitoring – it should match your calculation. - Check that the p95 latency stays flat during the test (no queuing delays).
- Search logs for “memory limit exceeded” or sandbox restarts.
- Confirm the effective setting:
gcloud run services describe SERVICE --format='value(spec.template.spec.containerConcurrency)'.
Key Limitations
- Maximum concurrency is 1,000 (hard limit).
- Min/max instance settings interact – minInstances=1 with high concurrency can waste money.
- CPU allocation (always vs. request‑only) affects background work but not concurrency limits.
- Traffic splitting creates independent scaling per revision – set concurrency on each revision separately.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.