Cloud Run CPU Allocation: Request-Based vs Always-Allocated Decision Guide
Cloud Run's CPU allocation mode changes cost, latency, and concurrency behavior at once. This guide compares request-based vs always-allocated CPU, shows the deploy flags, and gives a validation checklist.
23 Jul 2025, 04:03 UTC

The decision and why it matters
Every Cloud Run service revision runs in one of two CPU allocation modes. The default, request-based, gives your container CPU only while it is actively handling requests. The alternative, always-allocated (set with --cpu-always-allocated), keeps the vCPU provisioned for the entire lifetime of each instance, including idle time between requests.
This choice changes three things at once: what you pay, how your first request on a fresh instance behaves, and how concurrent requests share compute. Picking the wrong mode either wastes money on idle vCPU-seconds or leaves latency-sensitive traffic exposed to cold-start throttling.
Supported options compared
| Factor | Request-based (default) | Always-allocated |
|---|---|---|
| CPU billing | Only while processing requests | Entire instance lifetime |
| Idle instance CPU cost | None | Continuous (vCPU × uptime) |
| Background work between requests | Throttled/unreliable | Runs at full speed |
| Concurrency behavior | vCPU shared across concurrent requests | Full vCPU regardless of concurrency |
| Best fit | Bursty, low-duty-cycle HTTP APIs | CPU-bound, high-concurrency, or background-processing workloads |
Trade-offs in practice
Cost follows your duty cycle
The per-vCPU-second rate is the same in both modes; what changes is how many seconds you are billed for. If your service handles requests 10% of the time, request-based billing charges roughly a tenth of the CPU cost of always-allocated. As sustained utilization climbs past roughly 80%, the cost difference shrinks and always-allocated becomes easier to budget because the bill is a flat function of instance count. Measure your real duty cycle with the run/container/cpu/utilization metric in Cloud Monitoring before deciding — request-based shows a sawtooth pattern, always-allocated a flat line.
Latency and cold starts
A common misunderstanding: always-allocated CPU does not by itself eliminate cold starts. If your service scales to zero, the first request still waits for a container to start. What the mode changes is CPU availability after the instance exists. To get both warm instances and unthrottled CPU, combine --min-instances 1 with --cpu-always-allocated — at the cost of paying for that instance around the clock.
Note that startup probes, Secret Manager fetches, Cloud SQL connections, and VPC connector setup are network and auth operations; CPU allocation mode does not speed them up. If startup is slow because of initialization work, look at startup CPU boost and probe timeouts first.
Concurrency and background work
With request-based CPU, concurrent requests on one instance share the allocated vCPU, so CPU-bound work at high --concurrency can throttle. Always-allocated gives the instance its full vCPU regardless of request count. It is also the only supported way to reliably run background threads, scheduled in-process tasks, or connection keep-alives between requests, because request-based instances may be throttled to near-zero CPU when idle.
Implementation and validation
The mode is set per revision, so migration is a normal deploy with no downtime. Run these with the Cloud SDK installed and permission run.services.update on the service (typically via the Cloud Run Admin role):
# Deploy a revision with always-allocated CPU
gcloud run deploy mysvc \
--image=us-central1-docker.pkg.dev/PROJECT/REPO/mysvc:TAG \
--region=us-central1 \
--cpu=1 --memory=512Mi \
--min-instances=1 \
--cpu-always-allocated
# Confirm the revision configuration
gcloud run services describe mysvc --region=us-central1 \
--format='yaml(spec.template.spec.containers[0].resources)'Replace PROJECT, REPO, and TAG with your image coordinates. In the describe output, always-allocated revisions show the CPU allocation setting on the container; the absence of it indicates the request-based default.
To validate the effect under load, send traffic with a tool such as hey (hey -z 5m -c 10 https://YOUR_SERVICE_URL) and compare p50/p99 latency between revisions, then check the billing export for vCPU-second differences. Because this deploy creates a new revision, rollback is a traffic shift back to the previous revision with gcloud run services update-traffic — no rebuild needed.
Limitations to keep in mind
- Always-allocated reserves vCPU quota, not a dedicated physical core; steal time from host neighbors is rare on the Gen2 execution environment but possible.
- Pricing is region-dependent and changes over time. Verify current vCPU-second rates in Cloud Console Billing before committing to a cost model.
- Always-allocated with
--min-instances 0still scales to zero; you pay nothing while idle but keep the cold start on the first request. - Terraform users: the
google_cloud_run_v2_serviceresource exposes this as a CPU allocation field; check your provider version supports the v2 resource before planning.
Rule of thumb: start with the request-based default. Switch to always-allocated only when you have measured a concrete need — sustained high utilization, CPU-bound concurrency, or background work that must run between requests — and pair it with min-instances if cold starts are the actual problem.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.