High maxConcurrent vs Low maxConcurrent: Managing Latency and Contention in IBM Cloud Functions
27.5K reputation · 19 Jul 2026, 14:37 UTC
When configuring IBM Cloud Functions, the goal is to choose a concurrency setting that keeps response‑time latency low during bursts of concurrent invocations while avoiding CPU or memory throttling inside a shared container. The documented trade‑off is between raising maxConcurrent to allow multiple invocations to reuse a single warm instance (reducing cold‑start latency) and lowering maxConcurrent to isolate each invocation in its own container (eliminating contention but increasing cold‑start frequency). The decision depends on the expected peak concurrency, the function’s resource profile, and cost sensitivity.
What maxConcurrent value yields the best 99th‑percentile latency for a workload that regularly spikes to 50 concurrent invocations?
How does CPU utilization inside a shared container change as maxConcurrent increases from 1 to 10 for a function that consumes 200 MiB and 0.5 vCPU per invocation?
Under what concurrency threshold does the latency benefit of instance sharing become outweighed by throttling‑induced delays?