Answer
The overhead of signal‑based preemption grows roughly linearly with the number of active P structures (i.e., the effective GOMAXPROCS) when many goroutines are stuck in tight loops. Each P receives its own asynchronous preemption signal, so the total signal‑handling work scales with the count of Ps that are executing user code. The latency observed for any single goroutine is the sum of:
- time for the kernel to deliver the signal to the thread,
- entry into the runtime signal handler,
- checking the preemption flag and returning to the scheduler.
Under light load this totals a few‑to‑tens of microseconds on a typical Linux x86_64 system. When the run‑queue is saturated and many Ps are busy, kernel signal delivery can experience contention, pushing the tail latency higher.
Confirmed facts
- Asynchronous preemption was introduced in Go 1.14 using a platform‑specific signal (SIGURG on Unix/amd64).
- The mechanism does not depend on the loop body length; even an empty tight loop is preempted once the signal is processed.
- Measured latency on an idle system ranges from ~1 µs to ~30 µs.
- Disabling the feature with
GODEBUG=asyncpreemptoff=1 causes tight‑loop goroutines to run until they hit a safepoint.
Likely explanation
Each active P runs a dedicated OS thread that can receive the preemption signal independently. When the number of such threads approaches or exceeds the number of logical CPUs, the kernel must queue and deliver more signals per unit time, increasing the average delivery and handling time. This effect compounds with global run‑queue contention, leading to observable tail‑latency spikes in high‑concurrency workloads.
Steps to verify on your system
- Write a small program that spawns
N goroutines, each running an empty for {} loop.
- Run it with tracing enabled:
GOTRACEBACK=single go run -trace=trace.out prog.go.
- Examine the trace for
goroutine preemption events; note the timestamp delta between signal receipt and the goroutine’s state change.
- Repeat with different
N values (e.g., 1×, 2×, 4× GOMAXPROCS) and observe how the average and 99th‑percentile latency change.
- Optionally toggle the feature:
GODEBUG=asyncpreemptoff=0 (default) vs GODEBUG=asyncpreemptoff=1 and confirm that disabling it removes the preemption events and lets the loops run indefinitely.
Threshold guidance
No formal, version‑specific threshold is documented in the Go release notes. Empirical reports from the community suggest that when the number of busy Ps significantly exceeds the number of physical cores (e.g., > 1.5 × core count), the added signal‑handling cost can begin to outweigh the starvation‑prevention benefit, manifesting as increased tail latency. Whether this point is reached depends on your workload’s sensitivity to latency and the actual GOMAXPROCS setting.
Missing diagnostic detail: What is your current GOMAXPROCS value (or are you leaving it at the default)? Knowing this helps determine whether you are operating in the regime where signal overhead is likely to become noticeable.