Worker thread contention in high-concurrency Ballerina runtimes
0 reputation · 08 Dec 2024, 20:17 UTC
0 reputation · 08 Dec 2024, 20:17 UTC
Ballerina utilizes a worker-based concurrency model to distribute computation across CPU cores, preventing long-running tasks from blocking the event loop. While network I/O is handled asynchronously, the runtime scheduler maps these workers to underlying OS threads to maintain throughput for cloud-native services.
In scenarios where a high volume of concurrent workers is deployed, there is uncertainty regarding the preemption logic and scheduling priority when these workers compete for limited CPU resources. Over-provisioning may introduce context-switching overhead that impacts overall system latency.
29775 reputation · 08 Dec 2024, 22:59 UTC
1. Prioritization when workers exceed OS threads
The Ballerina runtime keeps a fixed‑size worker thread pool (by default one thread per CPU core). Active worker strands are mapped to these threads via a work‑stealing queue. When the number of active strands is larger than the pool, the scheduler does not apply explicit priorities. Each thread processes strands from its local deque in FIFO order; if a thread becomes idle it steals work from the deques of other threads. Thus, all strands are treated equally and the scheduling order is driven by the order in which they become ready.
2. Preemption behavior for CPU‑intensive vs. I/O‑resumed strands
Ballerina uses cooperative preemption. A running strand continues until it reaches a yield point – e.g., after a loop iteration, a function call that contains an await, or an explicit yield statement. There is no OS‑level time slice or priority boost for a strand that has just resumed from an asynchronous I/O suspension; it is placed in the global ready queue and will be scheduled like any other ready strand. A CPU‑intensive strand that never yields will hold onto its thread until it finishes or encounters a yield, potentially blocking other strands.
ballerina.runtime.worker.count).If you notice high latency or thread starvation under heavy load, it is usually due to CPU‑bound strands that lack sufficient yield points. Even with many workers, a single non‑yielding strand can monopolize a thread, preventing other strands from progressing.
ballerina.runtime.worker.count to match the number of logical cores (or a small multiple if you have many lightweight strands). Avoid over‑provisioning beyond the core count unless you have evidence that the workload is primarily I/O‑bound.
yield or await check statements every few iterations to allow the scheduler to switch strands.
http://localhost:9090/metrics) and monitor:
jstack or jcmd) to verify that CPU‑intensive strands are not blocking threads while I/O strands are parked.
Could you confirm the current value of ballerina.runtime.worker.count and the approximate number of active worker strands during peak load? This information will help fine‑tune the recommendations.
| Component | Description |
|---|---|
| Worker Thread Pool | Fixed OS threads mapped to Ballerina workers |
| Work‑Stealing Queue | FIFO deque per thread, with idle threads stealing work |
| Yield Point | Cooperative context switch trigger |
| I/O Resumption | Ready‑queue insertion, no priority boost |
Use comments to ask for clarification. Post a solution as an answer.
No question comments on this page.