Architecting Stateless HTTP Services on Cloud Run: A Minimalist Approach
Learn how to architect a minimalist, stateless HTTP service on Cloud Run, focusing on concurrency tuning, cold-start mitigation, and ingress trust boundaries.
21 Mar 2026, 17:58 UTC

The Problem: Balancing Scalability with Cold-Start Latency
When deploying containerized HTTP services, engineers often struggle to balance the cost-efficiency of scale-to-zero with the performance requirement of low-latency responses. The primary challenge in Cloud Run is managing the "Cold Start"—the delay incurred when the platform spins up a new container instance to handle an incoming request after a period of inactivity.
The takeaway for this architecture is to prioritize a small container footprint and precise concurrency tuning to minimize these delays while maintaining a strict stateless boundary.
The Smallest Suitable Design
For a standard stateless HTTP API, the most efficient architecture avoids unnecessary middleware. The minimal viable stack consists of:
- Container Image: A lightweight image (e.g., Alpine Linux or Distroless) containing the application binary and its runtime.
- Cloud Run Service: A deployment configured for request-based concurrency.
- Global External HTTP(S) Load Balancer: Used for custom domain mapping, SSL termination, and optional Cloud Armor integration for DDoS protection.
Configuration Example: Concurrency vs. Memory
A common failure point is mismatching the concurrency setting (the number of simultaneous requests one instance can handle) with the allocated memory. If concurrency is too high, the instance may hit an Out-of-Memory (OOM) error; if too low, you will trigger unnecessary cold starts.
# Example deployment command via gcloud CLI
# Run this from a terminal with Project Editor permissions
# Risk: Setting min-instances > 0 incurs ongoing costs
gcloud run deploy stateless-api \
--image gcr.io/[PROJECT_ID]/stateless-api:v1 \
--platform managed \
--region us-central1 \
--concurrency 80 \
--memory 512Mi \
--cpu 1 \
--min-instances 1 \
--max-instances 10
Trust and Data Boundaries
Cloud Run instances are ephemeral. Trust must be established at the ingress layer before the request reaches the container logic.
- Ingress Boundary: Use
--ingress internal-and-cloud-load-balancingto ensure the service is not exposed directly to the public internet, forcing all traffic through the Load Balancer. - Authentication Boundary: Implement IAM-based authentication for service-to-service communication. For public APIs, use an API Gateway or a dedicated authentication middleware within the container to validate JWTs (JSON Web Tokens).
- Data Boundary: The local filesystem is a temporary in-memory volume. Any data that must persist across requests must be stored in an external database (e.g., Cloud SQL) or object storage (e.g., Cloud Storage).
Operational Checks and Verification
To verify the health and scaling behavior of the deployment, monitor these specific metrics in the Google Cloud Console:
| Metric | What it Indicates | Target State |
|---|---|---|
| Container Instance Count | Scaling activity and cold start frequency | Stable during steady load; gradual climb during spikes |
| Request Latency (p99) | Impact of cold starts on end-users | Consistent latency regardless of request timing |
| 429 Error Rate | Saturation of max-instances |
Zero; indicates a need to increase instance limits |
Verification Step: To test the trust boundary, attempt to curl the run.app URL directly without an Authorization header. If the service is configured for "Allow unauthenticated invocations: No," the platform should return a 403 Forbidden before the request ever reaches your application code.
Failure Modes and Design Pivots
This minimalist design fails under specific conditions, requiring a shift in architecture:
- Stateful Requirements: If the application requires a persistent local disk for caching large files, Cloud Run is unsuitable. You must pivot to Google Kubernetes Engine (GKE) with Persistent Volume Claims (PVCs).
- Long-lived Connections: If the service requires WebSockets or non-HTTP TCP traffic that exceeds the request timeout limits, consider moving to a VM-based approach or GKE.
- Heavy Initialization: If the application framework (e.g., a large Spring Boot app) takes >10 seconds to start, the
min-instancessetting becomes mandatory to avoid unacceptable user latency.
Rollback Procedure
Because Cloud Run uses revisions, rolling back does not involve "undoing" a change but rather shifting traffic back to a known-good revision:
# Shift 100% of traffic back to a previous stable revision
gcloud run services update-traffic stateless-api --to-revisions [REVISION_NAME]=1000 replies
A thoughtful contribution can make all the difference. Be the first to share one.