Decide When to Disable Cloud Run CPU Throttling: Cost, Latency, and Best‑Practice Guide
Learn when to disable Cloud Run CPU throttling: cost vs latency, how to configure, monitor, and avoid common pitfalls. Includes gcloud and Terraform examples and practical monitoring tips.
18 Nov 2025, 10:55 UTC

When to Turn Off CPU Throttling on Cloud Run
Cloud Run’s default CPU allocation is request‑scoped: a vCPU is only active while a request is being processed. After the request finishes, the CPU is throttled to near‑zero, which saves money but adds a small CPU ramp‑up latency when the next request arrives. If your workload is latency‑sensitive—such as real‑time bidding, authentication, or high‑frequency trading—disabling throttling with --no-cpu-throttling (or cpu_throttling = false in YAML) keeps the full vCPU active at all times. This eliminates the ramp‑up penalty but incurs a linear cost for the wall‑clock time the instance is running.
How CPU Throttling Works in Practice
Below is a minimal Cloud Run v2 deployment that demonstrates the two options. The example uses a simple Go “Hello” service; the same principles apply to any language.
# Deploy with request‑scoped CPU (default)
# gcloud run deploy hello-default \
# --image gcr.io/cloudrun/hello \
# --region us-central1 \
# --cpu 1 \
# --memory 512Mi \
# --min-instances 1
# Deploy with always‑on CPU
# gcloud run deploy hello-nocpu \
# --image gcr.io/cloudrun/hello \
# --region us-central1 \
# --cpu 1 \
# --memory 512Mi \
# --min-instances 1 \
# --no-cpu-throttling
Both services share the same memory limit and minimum instance count, but the second keeps the vCPU running continuously. After deployment, you can use curl -w "%{time_total}\n" to measure the total round‑trip time for a request, and repeat the test after a five‑minute idle period to capture the cold‑start effect.
Cost Comparison
Using the Cloud Run pricing page as a reference, the cost for a 1‑vCPU / 512 MiB instance in us‑central1 is approximately $0.000024 per second when CPU‑throttling is disabled. With throttling enabled, you are charged only for the CPU time consumed by requests (minimum 100 ms per request). At 0 % utilization, the always‑on configuration will cost about $2.07 per month, while the throttled configuration will cost essentially nothing.
Monitoring to Verify CPU Behavior
run.googleapis.com/container/cpu/throttled_time– total time the CPU was throttled.run.googleapis.com/container/billable_instance_time– wall‑clock time billed for the instance.run.googleapis.com/request_latencies– request latency distribution; a spike in the 90th‑90th percentile after idle indicates throttling.
Set a Cloud Monitoring alert on container/cpu/throttled_time > 0 to detect unintended throttling, and on billable_instance_time to monitor cost spikes when you enable --no-cpu-throttling.
Common Mistakes and How to Avoid Them
- Enabling throttling for high‑concurrency workloads. Even if CPU is throttled, Cloud Run still allows the configured concurrency (default 80). If your requests are CPU‑bound, you may hit the concurrency limit and receive 503s. Consider lowering concurrency or disabling throttling.
- Relying on min‑instances alone.
--min-instanceskeeps memory allocated but does not prevent CPU throttling between requests. Combine it with--no-cpu-throttlingonly if you really need continuous CPU. - Ignoring quota limits. Always‑on CPU increases the number of CPU‑seconds you consume. If you have many services or high min‑instances, you may exceed the
CPU‑Seconds per monthquota. Request a quota increase before scaling. - Background threads stalling. Background goroutines or timers that run when no request is active will pause under throttling. If your service performs periodic health checks or cache invalidation, disable throttling or move that logic to a separate Cloud Scheduler job.
- Assuming throttling is the same across Cloud Run versions. Cloud Run v2 uses the same semantics as v1, but services migrated from v1 keep the original configuration. Verify the current setting in the console or via
gcloud run services describe.
Terraform Example
resource "google_cloud_run_v2_service" "example" {
name = "hello-nocpu"
location = "us-central1"
template {
containers {
image = "gcr.io/cloudrun/hello"
resources {
limits = {
cpu = "1"
memory = "512Mi"
}
}
# Disable CPU throttling
cpu_throttling = false
}
}
scaling {
min_instance_count = 1
}
}
When to Keep Throttling Enabled
- Low‑traffic, bursty workloads where idle periods are frequent and latency is not mission‑critical.
- Cost‑sensitive environments where you want to pay only for actual CPU usage.
- Services that perform significant background work that must not be paused.
When to Disable Throttling
- Real‑time or low‑latency APIs where p99 latency < 50 ms is required.
- Microservices that perform heavy CPU work per request (e.g., image processing, cryptographic operations).
- Environments with predictable, steady traffic where the cost difference is negligible compared to the latency benefit.
Takeaway
Disabling CPU throttling is a powerful tool for latency‑sensitive workloads, but it comes at a linear cost in wall‑clock time. Use --no-cpu-throttling only after measuring the latency benefit, verifying that your service can handle continuous CPU usage, and ensuring you have sufficient quota. Monitor cpu/throttled_time and billable_instance_time to catch unexpected behavior early.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.