Shipping to Cloud Run Without the Big Red Button: Traffic Splitting in Practice
Cloud Run runs multiple immutable revisions of one service and lets you split traffic between them per request. Here's a concrete canary workflow, plus the costs and limits nobody mentions.
02 Dec 2025, 14:20 UTC

Every deployment is a small bet. With Cloud Run, the bet doesn't have to be all-or-nothing: the platform lets you run multiple revisions of the same service side by side and decide, per request, which revision answers. Used well, this turns a risky cutover into a gradual dial you control — and a rollback into a traffic change instead of a redeploy.
This post walks through how revision traffic splitting actually works, a concrete canary rollout you can copy, and the costs and sharp edges that catch people out.
Revisions are immutable, and that's the point
Every deploy to a Cloud Run service creates a new revision — an immutable snapshot combining a container image with its configuration (concurrency, CPU, memory, timeout, environment). Immutability matters for two reasons:
- The previous revision keeps existing exactly as it was, so "rolling back" never means rebuilding anything.
- Two revisions can differ in configuration, not just code. You can canary a memory bump or a concurrency change the same way you canary a new image.
Traffic percentages are set at the service level and enforced by Cloud Run's routing layer per request. Percentages must sum to 100, and propagation is eventually consistent — a split you just applied may take a short while to be visible everywhere, so don't judge it from the first two requests.
A worked canary rollout
Assume a service payments in us-central1, currently serving revision payments-00042. You want to ship a new image safely. Run these from a machine with the gcloud CLI authenticated, using an identity with the Cloud Run Developer (or Admin) role on the service.
1. Deploy the new revision with no traffic. The --no-traffic flag is the key move — the revision exists and is health-checked, but no user hits it:
gcloud run deploy payments \
--image=us-central1-docker.pkg.dev/PROJECT/repo/payments:v1.7.0 \
--region=us-central1 \
--no-traffic --tag=canaryThe --tag gives the revision its own URL of the form https://canary---payments-<hash>.run.app. Hit that URL directly to smoke-test the new revision before it receives any shared traffic. A useful trick: have your app return the revision name (available via the K_REVISION environment variable) in a response header, so you can confirm which revision answered any request.
2. Shift a small slice of traffic.
gcloud run services update-traffic payments \
--region=us-central1 \
--to-tags=canary=10Now roughly 10% of requests land on the new revision. Watch the Cloud Run metrics dashboard (or your own monitoring) for request count, error rate, and latency, filtered by revision — the per-revision breakdown is what makes this observable.
3. Promote or retreat. If metrics look healthy, step up to 50%, then 100%:
gcloud run services update-traffic payments \
--region=us-central1 \
--to-tags=canary=100If error rate climbs, rollback is a traffic command, not a redeploy:
gcloud run services update-traffic payments \
--region=us-central1 \
--to-revisions=payments-00042=100Expected checks: after each step, gcloud run services describe payments --region=us-central1 shows the traffic allocation, and your response-header trick confirms routing. Verify rollback by watching the error rate return to baseline within a few minutes.
Warming the canary — and what it costs
A revision at 10% traffic still scales from zero by default, so its first users may eat cold-start latency. The common fix is setting --min-instances=1 (or more) on the new revision so a warm instance is always ready. Two cautions:
- Min instances bill continuously, even at 0% traffic. For latency-tolerant internal tools, they're often wasted money; for user-facing canaries, they're usually worth it.
- CPU allocation behavior is version- and configuration-sensitive. With the default (CPU allocated only during request processing), idle min instances cost less than with always-allocated CPU — but always-allocated CPU changes cold-start and background-processing behavior too. Check the current docs for your configuration before assuming the cost model.
Limitations worth knowing before you rely on this
- No request affinity. Splitting is per request, not per user. The same user can bounce between revisions on consecutive calls. If revisions behave differently in ways a session depends on, you need an external session store (or to avoid splitting across incompatible versions).
- Two revisions means two of everything in your logs. Always filter logs and metrics by revision name, or you'll average away the very regression you're looking for.
- Revision accumulation. Old revisions stick around. Unused ones serving 0% traffic don't serve requests, but they add clutter and configuration-drift risk. Prune them on a schedule.
- Concurrency caps still apply. Traffic splitting doesn't raise your per-instance concurrency limit; a hot revision still scales out the same way.
The takeaway
The pattern to adopt: deploy with --no-traffic and a tag, smoke-test the tagged URL, shift 10%, watch per-revision error rate and latency, then promote or revert with a single traffic command. Your actionable next step is the smallest one — add the K_REVISION response header to your app today, so that when you do your first split, you can actually see it working.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.