Short answer
There is no single maxConcurrent value that guarantees the best p99 latency for a 50‑concurrent‑invocation workload — the right setting depends on how much CPU and memory your function actually uses per activation. That said, the documented trade‑off gives a defensible starting point: for a function that needs roughly 0.5 vCPU and 200 MiB per invocation, a maxConcurrent of 2–4 per container is a reasonable initial value, followed by a load test to find the real ceiling. Setting it to 10 would stack ~5 vCPU of demand into one container and almost certainly cause CPU throttling and worse tail latency than the cold starts you were trying to avoid.
What is confirmed vs. what is workload‑dependent
Confirmed behavior: maxConcurrent controls how many activations may run inside a single warm container. Higher values reuse warm instances (fewer cold starts, less queueing) but make activations share the container's CPU, memory, and file descriptors. Lower values isolate each invocation but force the platform to spin up more instances, so bursts hit cold starts and queueing delay more often. If you push concurrency past platform limits, activations can be rejected with 429‑style throttling errors.
Workload‑dependent (not answerable without measurement): the exact p99‑optimal value and the exact threshold where sharing stops helping. These depend on whether your function is CPU‑bound, I/O‑bound (waiting on a database or API), or memory‑bound. An I/O‑bound function can tolerate much higher maxConcurrent than a CPU‑bound one because waiting threads don't consume cores.
CPU utilization as maxConcurrent rises from 1 to 10
For a CPU‑bound function consuming ~0.5 vCPU per invocation, demand scales roughly linearly: maxConcurrent 1 ≈ 0.5 vCPU, 2 ≈ 1 vCPU, 4 ≈ 2 vCPU, 10 ≈ 5 vCPU requested inside one container. Once summed demand exceeds the CPU actually allocated to the container, the kernel scheduler (and cgroup CPU limits) start time‑slicing or throttling the activations. Each invocation then runs slower, and because all co‑resident invocations degrade together, tail latency (p99) rises sharply even if average latency looks acceptable. Memory is the harder wall: 10 × 200 MiB = ~2 GiB, which can exceed the container memory limit and trigger out‑of‑memory failures rather than graceful slowdown.
If your function is mostly I/O‑bound, real CPU usage per invocation may be far below 0.5 vCPU, and higher maxConcurrent can be safe — which is exactly why you must measure rather than assume.
Finding the threshold where sharing stops paying off
The crossover point is where the marginal cold start you avoid is cheaper than the marginal throttling delay you introduce. Find it empirically:
- Check the current setting:
ibmcloud fn action get <actionName>. - In a staging namespace, run a load test (k6 or Apache Bench) that sustains 50 concurrent requests, repeating for maxConcurrent = 1, 2, 4, 8.
- Record p50/p99 latency, error rate, and activation queue length from the IBM Cloud Functions monitoring dashboard for each run.
- The suitable ceiling is the highest value where p99 is still flat and errors are zero. When p99 bends upward or throttling/OOM errors appear, back off one step.
Practical guidance for your case
For 50 concurrent invocations of a 0.5 vCPU / 200 MiB function: with maxConcurrent = 2, the platform needs ~25 warm instances at peak — cold starts only while scaling up, minimal contention. With maxConcurrent = 10, you need only 5 instances, but each is oversubscribed on CPU, so steady‑state p99 will likely be worse. Start at 2–4, load test, and adjust. If the function turns out to be I/O‑heavy (e.g., waiting on Db2 or an external API), test higher values like 8–10 as well, since idle waiting doesn't contend for CPU.
One caveat: these resource figures assume the 0.5 vCPU / 200 MiB numbers reflect actual usage under load, not just the configured action memory limit. If you haven't profiled real per‑invocation CPU consumption, that measurement is the one missing detail that would change the recommendation.