Managing Cold Starts in Cloud Run with Min Instances, Concurrency, and CPU Allocation
Learn how to reduce Cloud Run cold‑start latency with min‑instances, tune concurrency, and choose CPU allocation, while balancing cost and throughput for production APIs.
16 Jun 2026, 02:30 UTC

Problem: The Cold‑Start Pain Point
When a Cloud Run service is idle, Google automatically scales all containers to zero. The first request that arrives after this idle period forces a new container to start, which can add 200–400 ms of latency (sometimes more for large images or complex init code). For APIs that serve latency‑sensitive clients—mobile apps, real‑time dashboards, or edge‑compliant microservices—this tail latency can push p99 latency above SLA targets.
Thesis: Use Min‑Instances, Concurrency, and CPU Allocation Wisely
Cloud Run gives three knobs that directly influence cold‑start behavior and cost:
- Min‑instances – keeps a warm pool of containers available.
- Concurrency per instance – determines how many requests a single container can handle at once.
- CPU allocation mode – decides when CPU is reserved for a container.
By configuring these knobs together, you can keep the p99 latency low for bursty traffic without over‑provisioning resources that drive up the bill.
Section 1: Min‑Instances – The Warm Pool
Setting --min-instances=1 (or higher) tells Cloud Run to keep at least that many containers running at all times, even when traffic is zero. These instances will be immediately ready to serve the next request, eliminating the cold‑start delay for the first few hits.
However, you pay for CPU and memory on those warm containers continuously, regardless of load. For a low‑traffic service that runs 24/7, the cost of a single warm instance can be noticeable.
Example Configuration
gcloud run deploy my-api \
--image gcr.io/my-project/my-api:latest \
--region us-central1 \
--min-instances 1 \
--max-instances 10 \
--concurrency 80 \
--cpu=always \
--platform managed
Replace my-project and my-api with your own values. The command above:
- Deploys the image to Cloud Run in the
us-central1region. - Keeps one container warm at all times.
- Allows up to ten containers to scale out when traffic spikes.
- Sets a concurrency of 80 requests per container.
- Allocates CPU continuously, not just during request processing.
Section 2: Concurrency – Balancing Throughput and Latency
Concurrency is the number of requests a single container can handle simultaneously. The default is 80, but you can raise it to 1000 if your container is lightweight and CPU‑bound. Increasing concurrency reduces the number of containers needed for a given request rate, which in turn lowers the warm‑pool cost. But it also increases the queue length when the container is saturated, raising tail latency.
When traffic is bursty, a higher concurrency can keep the warm pool effective: the single warm instance can absorb a spike of 80–100 requests without needing to spin up a new container.
Section 3: CPU Allocation – CPU‑Always vs. CPU‑On‑Demand
Cloud Run offers two CPU allocation modes:
--cpu=always– CPU is reserved for the container 100 % of the time. This is ideal for CPU‑heavy workloads, ensuring consistent performance even when the container is idle.--cpu=on-demand– CPU is allocated only during request handling. This saves cost when the container is idle but can introduce a delay when the first request arrives after a period of inactivity.
For latency‑sensitive services, --cpu=always is usually preferred because it removes the “CPU‑wait” part of the cold start. The trade‑off is a modest increase in cost for idle containers.
Section 4: Trade‑Offs and Practical Checks
Below is a quick cost‑benefit comparison for a typical API that sees 1 k requests per minute during peak hours and 50 requests per minute during off‑peak:
| Configuration | Avg. CPU (cores) | Memory (GiB) | Estimated Monthly Cost (USD) |
|---|---|---|---|
| Min‑instances 0, Concurrency 80, CPU on‑Demand | 0.05 | 0.5 | $30–$40 |
| Min‑instances 1, Concurrency 80, CPU always | 0.1 | 0.5 | $70–$80 |
| Min‑instances 1, Concurrency 1000, CPU always | 0.1 | 0.5 | $70–$80 |
These numbers are rough estimates based on the Cloud Run pricing model (CPU at $0.000024 per vCPU‑second, memory at $0.000004 per GiB‑second). The key takeaway: keeping one instance warm almost doubles the cost compared to scaling to zero, but it can shave 200–300 ms off the p99 latency for bursty traffic.
To validate your configuration:
- Use
gcloud beta run services describe my-api --region us-central1to confirm theminInstancesvalue. - Check Cloud Monitoring metrics:
run.googleapis.com/container_cpu_utilizationandrun.googleapis.com/container_memory_utilizationshould stay low when idle. - Run a controlled load test with a tool like k6 or wrk to capture p50/p99 latency under warm and cold conditions.
Actionable Closing: Pick the Right Knobs for Your Workload
1. Measure your traffic pattern. If you have predictable spikes or a latency SLA, consider a warm pool.
2. Start with min‑instances 1. Observe cost and latency; if the cost is too high, raise concurrency or lower max instances.
3. Choose CPU allocation based on workload. CPU‑heavy services benefit from --cpu=always even if it costs a bit more.
4. Iteratively test. Use Cloud Monitoring and load testing to confirm that the chosen settings meet your latency goals without over‑provisioning.
By balancing min‑instances, concurrency, and CPU allocation, you can keep Cloud Run services responsive under bursty traffic while staying within budget.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.