Architecting Stateless HTTP Services on Cloud Run: A Minimalist Approach
Learn how to architect a minimalist, stateless HTTP service on Cloud Run, focusing on concurrency tuning, cold-start mitigation, and ingress trust boundaries.
ReadMeFeed / Community knowledge
Real questions. Useful conversations. Find the people who know your stack.
Learn how to architect a minimalist, stateless HTTP service on Cloud Run, focusing on concurrency tuning, cold-start mitigation, and ingress trust boundaries.
Cloud Run's request concurrency setting trades cost against latency. Learn when to set it high, low, or 1, with a worked example and a load-test verification plan.
Learn how to set Cloud Run concurrency and CPU allocation to keep a user‑facing API under 200 ms latency while controlling cost. A quick decision table, trade‑off analysis, and deployment example are included.
Goal Enable precise, non‑integer traffic allocation to revisions for canary and blue‑green deployments, such as 33.3%/33.3%/33.4% splits. Constraints The current Cloud Run Traffic Splitting API accepts only whole‑percentage values (e.g., 25%, 75%). Both the gcloud CLI and Cloud Console reject decimal inputs with validation errors. No public roadmap or releas
The goal is to perform a zero‑downtime migration of a small application by deploying a new Cloud Run revision and gradually shifting traffic via traffic splitting, keeping the service URL stable. Because Cloud Run does not provide session affinity and treats each request independently, long‑lived connections (e.g., WebSocket, persistent HTTP, or database poo