Diagnosing Cloud Run Latency and Timeout Failures: A Step-by-Step Guide
A diagnostic guide for Cloud Run latency and timeout failures: recognize the condition, run ordered checks (logs, concurrency, resources, timeout, revision), apply targeted fixes, and know when to escalate. Includes gcloud commands, verification steps, and a concrete rollback example.
01 Jan 2026, 14:24 UTC

The Problem: Unexplained 504s and Spiking Latency
Your Cloud Run service returns HTTP 504 Gateway Timeout errors, or request latency jumps from tens of milliseconds to 30+ seconds without a code change. The symptom is visible in Cloud Monitoring dashboards and user-facing error rates, but the root cause is often a configuration mismatch rather than application logic. This guide walks through a repeatable diagnostic process to isolate the bottleneck—whether it's cold starts, concurrency saturation, resource limits, or timeout misconfiguration—and apply the minimal fix.
Recognizable Condition
- HTTP 504 responses increasing on a managed Cloud Run service
- Request latency percentile (p95/p99) climbing above the configured request timeout
- No corresponding deployment or code change at the time of onset
- Cloud Run revision history shows a stable revision receiving traffic
Cause and Diagnostic Quick-Reference Table
| Observed Symptom | Likely Cause | Primary Diagnostic Signal |
|---|---|---|
| 504 on first request after idle period | Cold start exceeds request timeout | Log entry: "Container started in X.XXs" near timeout value |
| 504 under sustained load | Concurrency limit reached; requests queued then timed out | Metric: run/container/request_count plateau at concurrency setting |
| OOM kills or CPU throttling in logs | Memory/CPU allocation too low for workload | Log: "Memory limit exceeded" or container/cpu/utilization near 100% |
| Intermittent 504 with no resource pressure | Request timeout set lower than actual processing time | Compare request_timeout setting vs. observed latency histogram |
| Errors spike after new revision rollout | Regression in startup or request handling | Revision traffic split shows new revision receiving errors |
Ordered Diagnostic Checks
Run these checks in sequence. Each check either isolates the cause or rules out a category.
1. Examine Cloud Run Logs for Startup or Crash Messages
Open the Logs Explorer in the Google Cloud Console for your service (resource.type="cloud_run_revision"). Filter for severity >= ERROR and look for:
- Container startup duration messages (e.g., "Container started in 8.2s")
- OOM kill events: "Memory limit exceeded" or exit code 137
- Application-level stack traces on request handling
- Health check failures if using custom health endpoints
Where to run: Cloud Console > Logging > Logs Explorer, or gcloud logging read with appropriate filter.
Permissions: logging.logEntries.list on the project.
Expected check: If startup time approaches the request timeout (default 300s, often lowered), cold start is the culprit. If OOM appears, move to check 3.
2. Verify Configured Concurrency vs. Actual Request Count
Cloud Run defaults to 80 concurrent requests per instance (configurable up to 1000). When incoming requests exceed concurrency × active_instances, excess requests queue until the request timeout expires.
gcloud run services describe SERVICE_NAME --platform managed --region REGION \
--format="value(template.spec.containerConcurrency)"
Then check the metric run/container/request_count grouped by revision_name in Cloud Monitoring. If the per-instance request count flattens at the concurrency value while latency rises, you've hit the concurrency ceiling.
Risk: Raising concurrency beyond your application's thread-safe capacity can cause race conditions, connection pool exhaustion, or degraded per-request latency. Test with a load generator before increasing in production.
3. Compare Memory/CPU Usage to Allocated Limits
In Cloud Monitoring, chart run/container/memory/utilization and run/container/cpu/utilization for the affected revision. Sustained utilization above 80% indicates resource pressure.
- Memory: If utilization trends to 100% before OOM kills, increase the memory limit (e.g.,
--memory=2Gi). - CPU: For CPU-bound workloads, enable "CPU always allocated" (
--cpu-always-allocatedorcpu.alwaysAllocated: truein YAML) so the instance retains full CPU during idle periods, eliminating throttle-on-wake latency.
Cost note: Raising memory or CPU increases per-instance cost. Monitor billing after changes.
4. Confirm Request Timeout Setting
The request timeout defaults to 300 seconds but is often reduced. If your p99 latency exceeds this value, requests are terminated at the platform level before your code completes.
gcloud run services describe SERVICE_NAME --platform managed --region REGION \
--format="value(template.spec.timeoutSeconds)"
Compare this value to the latency histogram in Monitoring. If timeout < p99 latency, increase it (maximum 3600s) or optimize the slow code path.
5. Review Active Revision and Traffic Split
A recent deployment may have introduced a regression. Check which revisions receive traffic:
gcloud run services describe SERVICE_NAME --platform managed --region REGION \
--format="table(status.traffic[].revisionName,status.traffic[].percent,status.traffic[].latestRevision)"
If a non-latest revision shows errors while the latest is healthy (or vice versa), roll traffic to the stable revision while investigating.
Fixes Tied to Findings
| Finding | Fix | Command / Action |
|---|---|---|
| Cold start > request timeout | Reduce initialization time; enable CPU always allocated; consider background initialization | gcloud run services update SERVICE --cpu-always-allocated --region REGION |
| Concurrency limit saturated | Increase concurrency if app is thread-safe; otherwise increase max instances | gcloud run services update SERVICE --concurrency=200 --max-instances=50 --region REGION |
| Memory limit exceeded | Increase memory allocation | gcloud run services update SERVICE --memory=2Gi --region REGION |
| CPU throttling on wake | Enable CPU always allocated | gcloud run services update SERVICE --cpu-always-allocated --region REGION |
| Request timeout too low | Raise timeout to exceed p99 latency + margin | gcloud run services update SERVICE --timeout=600s --region REGION |
| Recent revision regression | Rollback traffic to prior revision | gcloud run services update-traffic SERVICE --to-revisions=PREV_REVISION=100 --region REGION |
Concrete Example: Diagnosing a 504 Spike After a Dependency Upgrade
A team deploys a new revision that upgrades a database driver. Within minutes, 504 errors rise to 12% of requests. Diagnostic steps:
- Logs: No OOM, no startup errors. Container starts in 1.2s (timeout is 60s).
- Concurrency: Service configured at default 80. Metric shows 78–80 concurrent requests per instance during the spike.
- Resources: Memory at 45%, CPU at 60%—no pressure.
- Timeout: Configured at 60s. Latency histogram shows p99 at 58s, p99.9 at 62s.
- Revision: New revision receiving 100% traffic.
Root cause: The new driver introduces a connection pool contention bug under load, causing individual requests to occasionally exceed 60s. The fix path:
- Immediate: Roll back to previous revision (
gcloud run services update-traffic ...). - Short-term: Increase timeout to 120s while debugging the driver issue.
- Long-term: Fix connection pool sizing in application code.
Verification After Applying a Fix
Deploy a minimal test service with the same configuration changes, then run a load test:
# Deploy test service (assumes container image exists)
gcloud run deploy test-latency-fix --image=REGION-docker.pkg.dev/PROJECT/REPO/IMAGE \
--platform managed --region REGION --concurrency=200 --memory=2Gi --timeout=600s
# Load test from a separate machine or Cloud Shell
hey -z 2m -c 50 https://test-latency-fix-REGION.run.app/health
Observe in Cloud Monitoring:
- Request latency p95/p99 should drop below the new timeout
- Error rate (5xx) should approach 0%
- Active instance count should scale smoothly without hitting max-instances
Compare the same metrics before and after the fix on the production service to confirm improvement.
Escalation Criteria
Engage Google Cloud Support when:
- 5xx errors persist after applying all relevant configuration fixes above
- Service hits the
max-instanceslimit unexpectedly and cannot scale further (quotas already increased) - Advanced networking is required (VPC egress controls, Private Service Connect, custom DNS) and configuration changes don't resolve connectivity-related timeouts
- You observe platform-level anomalies: repeated "instance evacuation" events, unexplained revision health check failures, or metrics gaps in Monitoring
Limitations and Practical Checks
- This guide covers managed Cloud Run (fully managed). Cloud Run for Anthos has different scaling knobs and networking paths.
- Concurrency behavior assumes the application handles concurrent requests correctly. If your code uses global mutable state, increasing concurrency will cause data corruption before it solves latency.
- Cold start optimization (min instances, CPU always allocated, background initialization) trades cost for latency. Set
--min-instancesonly for latency-sensitive paths; each idle instance incurs full cost. - Always verify changes in a non-production environment first. The
heyload test above is a starting point, not a substitute for realistic traffic patterns.
Quick Reference: Key gcloud Commands
| Task | Command |
|---|---|
| View current service config | gcloud run services describe SERVICE --platform managed --region REGION |
| Update concurrency and max instances | gcloud run services update SERVICE --concurrency=VALUE --max-instances=VALUE --region REGION |
| Update memory/CPU | gcloud run services update SERVICE --memory=2Gi --cpu=2 --region REGION |
| Enable CPU always allocated | gcloud run services update SERVICE --cpu-always-allocated --region REGION |
| Update request timeout | gcloud run services update SERVICE --timeout=600s --region REGION |
| Rollback traffic | gcloud run services update-traffic SERVICE --to-revisions=REVISION=100 --region REGION |
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.