Diagnosing and Resolving Request Failures in Cloud Run
A diagnostic guide for resolving HTTP 429, 502, and timeout errors in Cloud Run by tuning concurrency, min-instances, and IAM permissions.
27 Jun 2026, 19:20 UTC

Identifying the Root Cause of Cloud Run Failures
When a Cloud Run service begins failing under load or exhibiting erratic latency, the cause is rarely a generic "crash." Most issues stem from a mismatch between the service configuration (concurrency, timeouts, and instance counts) and the actual traffic patterns of the application.
Diagnostic Matrix: Symptoms to Causes
| Symptom | Primary Suspect | Key Metric to Check |
|---|---|---|
| HTTP 429 (Too Many Requests) | Max Concurrency Limit | Container Instance Count vs. QPS |
| Sporadic Latency / HTTP 502 | Cold Starts | Instance Start-up Time |
| Deadline Exceeded / Timeout | Request Timeout Setting | Request Duration (Logs) |
| Revision Creation Failed | IAM Permissions | Service Account Roles |
Step-by-Step Resolution Path
1. Resolving HTTP 429 (Concurrency Limits)
Cloud Run limits how many concurrent requests a single container instance can handle. If your max-concurrency is set too low, Cloud Run will attempt to spin up new instances. If the scaling limit is hit or the request rate spikes faster than instances can start, you will see 429 errors even if CPU usage is low.
- Check: Compare your current
max-concurrencysetting against your peak requests per second (QPS). - Fix: Increase the concurrency limit. For most web applications, a value between 80 and 100 is a stable starting point, provided the application is asynchronous.
- Command: Run this via gcloud CLI (requires
roles/run.admin):gcloud run services update [SERVICE_NAME] --concurrency 80 --region [REGION] - Risk: Increasing concurrency without increasing memory/CPU can lead to Out-of-Memory (OOM) kills if each request consumes significant RAM.
2. Eliminating Cold Start Latency
Cold starts occur when Cloud Run scales from zero instances to one. This causes a significant delay for the first request, which can trigger 502 errors in sensitive client timeouts.
- Check: Review Cloud Monitoring for "Instance start-up time." If spikes correlate with periods of inactivity, you have a cold start problem.
- Fix: Set a minimum number of instances to keep containers "warm."
- Command:
gcloud run services update [SERVICE_NAME] --min-instances 1 --region [REGION] - Risk: Setting
--min-instancesincurs ongoing costs as you are paying for the allocated resources even when no traffic is present.
3. Fixing Request Timeouts
By default, Cloud Run terminates requests that exceed 60 seconds. Long-running tasks, such as large PDF generation or complex database migrations, will return a "deadline exceeded" error.
- Check: Search logs for
deadline exceededand check the duration of the failed requests. - Fix: Increase the timeout limit (up to 600 seconds).
- Command:
gcloud run services update [SERVICE_NAME] --timeout 300 --region [REGION] - Alternative: If the task takes longer than 10 minutes, move the logic to an asynchronous queue (e.g., Cloud Pub/Sub) and return a 202 Accepted response to the client.
4. Troubleshooting Deployment Failures
If a deployment fails with Revision creation failed: container image not found, it is usually an identity issue rather than a missing image.
- Check: Verify the service account used by Cloud Run has the
roles/artifactregistry.readerrole on the specific Artifact Registry repository. - Fix: Grant the permission via IAM:
gcloud projects add-iam-policy-binding [PROJECT_ID] --member="serviceAccount:[SERVICE_ACCOUNT_EMAIL]" --role="roles/artifactregistry.reader"
Verification and Escalation
To verify these fixes, use a load-testing tool (such as hey) to send a steady stream of requests. Monitor the Cloud Run Metrics tab to ensure that the 429 error rate drops to zero and that the instance count stabilizes.
When to Escalate
If you have applied these fixes and still experience issues, escalate your investigation to the following areas:
- High Latency after Concurrency Increase: Profile your application for blocking I/O or check for database connection pool exhaustion.
- Persistent 502s: Check the health of your Serverless VPC Access connector if you are connecting to a private database.
- High Egress Costs: If outbound traffic costs are spiking, ensure you are using a VPC Connector to route internal traffic rather than traversing the public internet.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.