Answer the question
For a service that receives only a handful of requests per day, the first setting to adjust is minimum instances. Setting min‑instances to 0 ensures you aren’t billed for idle instances during the long periods of no traffic.
Effect on cold‑start latency
With min‑instances=0, Cloud Run will spin up a new instance only when a request arrives. This increases cold‑start latency compared to keeping a warm instance, but for sporadic traffic the trade‑off is usually acceptable. You can mitigate the impact by keeping the default CPU allocation mode (“CPU only when request”) and by using a higher concurrency (default 80) to keep fewer instances alive.
Compatibility constraints
- All workloads supported by Cloud Run can set
min‑instances to 0.
- If you rely on secret‑volume mounts that require a running instance, they will still work because the instance is created on first request.
- CPU‑only‑when‑request mode is the default; if you have a legacy
always‑on‑CPU setting, switch it back to the default to avoid paying for idle CPU.
Likely explanation (based on best‑practice guidance)
After setting min‑instances=0, you can further tweak:
- Concurrency: Increase to 80 (default) or higher if your container is stateless and can handle many simultaneous requests. This reduces the number of instances needed.
- Memory: Lower to the smallest size that satisfies peak usage; each memory tier has a fixed per‑instance cost.
- Choose a region with lower compute or egress prices if traffic is not region‑sensitive.
Confirmed facts (from GCP documentation and cost‑optimization guidance)
- Setting
min‑instances=0 removes the idle‑instance charge.
- Cold‑start latency is proportional to the time Cloud Run takes to pull the image and start the container.
- CPU allocation mode can be set to
CPU only when request (default) or CPU always; the former saves cost when idle.
- Concurrency limits are per instance; higher concurrency reduces the total number of instances.
- Memory tiers are billed per instance; reducing memory lowers the per‑instance rate.
Next steps for your environment
- Set
min‑instances to 0 in the Cloud Run console or via gcloud run services update SERVICE --min-instances=0.
- Verify the
CPU allocation mode is CPU only when request (default).
- Monitor
instance count and latency in Cloud Monitoring after the change.
- If latency spikes, consider raising
concurrency (e.g., to 80) or lowering memory slightly, then re‑monitor.
- Check Cloud Billing for the cost difference before and after the change.
Missing diagnostic detail
To fine‑tune concurrency and memory, could you confirm whether the workload is more CPU‑bound or memory‑bound during peak requests? This will influence whether you should increase concurrency or reduce memory.