Designing Around k6's Virtual User Model: When One Generator Is Enough
k6 runs each Virtual User in its own Goja runtime on a goroutine — great for I/O-bound load, dangerous for CPU-heavy scripts. An architecture note on sizing one generator correctly and knowing when to split.
04 May 2026, 05:00 UTC

Most k6 load tests fail for the wrong reason: the generator saturates before the target does, and the resulting latency numbers describe your laptop, not your service. The fix is not a bigger machine by default — it is understanding how k6 actually executes Virtual Users (VUs), so you can size the smallest design that keeps the generator out of the results.
Requirements
A load test design has to answer three questions before any script is written:
- What load shape? Steady request rate, ramping concurrency, or a fixed number of iterations.
- What is the pass/fail line? Thresholds on latency, error rate, or checks.
- How will you know the generator was not the bottleneck? Without this, every other number is suspect.
Everything else — distribution, cloud execution, multiple scenarios — is a response to measured limits, not a starting point.
How k6 actually runs VUs
k6 is a Go binary that embeds Goja, a JavaScript runtime written in Go. Two properties of this architecture drive the whole design:
- Each VU gets its own isolated JS runtime. There is no shared mutable state between VUs at the script level. A variable you set in one VU's iteration is invisible to every other VU. This is a feature: it eliminates an entire class of race conditions, but it also means you cannot build coordination logic in script globals.
- Each VU runs in a goroutine, not an OS thread. The Go scheduler multiplexes thousands of goroutines onto a small pool of threads. This is why a single k6 process can plausibly drive thousands of concurrent VUs — as long as the VUs spend most of their time waiting on network I/O rather than computing.
The second point has a sharp edge. Goja is an interpreter, not a JIT; it is considerably slower than V8 for CPU-heavy work. If your script does cryptographic signing, large JSON transformations, or heavy string processing per iteration, VU CPU cost stops being negligible and the generator becomes the system under test. Measure before assuming.
The smallest suitable design
For a steady-state API test, the smallest defensible design is one k6 process, one scenario, and the constant-arrival-rate executor. Arrival-rate executors schedule iterations per unit time and start VUs as needed, which models real traffic better than fixing a VU count and letting response time dictate throughput (the coordinated omission problem).
import http from 'k6/http';
import { check } from 'k6';
import { SharedArray } from 'k6/data';
const users = new SharedArray('users', function () {
return JSON.parse(open('./users.json'));
});
export const options = {
scenarios: {
steady: {
executor: 'constant-arrival-rate',
rate: 200, // iterations per second
timeUnit: '1s',
duration: '10m',
preAllocatedVUs: 50, // pool sized from a smoke test, not a guess
maxVUs: 400,
},
},
thresholds: {
http_req_failed: ['rate<0.01'],
http_req_duration: ['p(95)<500'],
},
};
export default function () {
const u = users[__VU % users.length];
const res = http.get(`https://api.example.com/orders?user=${u.id}`);
check(res, { 'status 200': (r) => r.status === 200 });
}Run this with k6 run test.js from a shell on the generator machine; no elevated privileges are needed, but the process must be allowed enough open file descriptors (check ulimit -n, since each in-flight connection consumes one). The meaningful placeholders are the target rate, the host, and the fixture file. The risk to note: if the target rate exceeds what the generator can produce, k6 logs dropped iterations — treat any dropped iteration as invalidating the run, not as a target failure.
Trust and data boundaries
The script executes entirely on the generator with two notable capabilities: outbound network access and file reads in the init context (the top-level code that runs once per VU at startup). That means credentials and fixtures live on the generator and must be protected accordingly — and conversely, nothing sensitive should be pasted into a script that gets committed.
Because VUs cannot share mutable state, read-only data flows in through exactly two channels: SharedArray for files (parsed once, shared read-only across VUs, which keeps memory flat) and __ENV for environment variables. If you find yourself wanting a shared counter or queue between VUs, that data belongs in the target system or an external service, not in the script.
Operational checks during the run
- Generator CPU and memory. Watch the generator, not just the dashboard. Sustained CPU near saturation means latency samples are inflated by scheduling delay.
- Open file descriptors and ephemeral ports. High connection churn can exhaust the ephemeral port range; this shows up as connection errors that look identical to target failures in the summary.
- Dropped iterations. For arrival-rate executors, non-zero dropped iterations mean
maxVUsor generator capacity was insufficient. - Thresholds as the verdict. The end-of-test summary plus thresholds give you pass/fail; compare against server-side metrics (CPU, queue depth, p95 from the server's own instrumentation) to confirm which side degraded.
Failure modes
- Generator exhaustion. CPU, memory, file descriptors, or ephemeral ports run out; latency and error metrics rise on the client side while the server is idle. This is the most common way load tests lie.
- Script exceptions. An uncaught exception aborts that VU's current iteration and appears in the log and summary — it does not fail the target, but it silently reduces applied load.
- CPU-bound script logic. Heavy per-iteration computation under Goja caps throughput regardless of VU count.
When the design has to change
Move beyond one process only when measurement says so: generator CPU saturated while the target is healthy, or dropped iterations at your target rate. Then the options are splitting scenarios across multiple k6 instances (the k6 Kubernetes operator coordinates this) and aggregating results externally, or using managed cloud execution for geographically distributed load. A useful validation before trusting any single-process ceiling: run the same total load split across two instances and compare aggregate throughput — if the split run delivers more, your single-process number was a generator limit.
Executor semantics and module APIs shift between releases, so confirm behavior against the docs for the version k6 version reports before relying on specifics. And never quote per-VU memory or CPU overhead you have not measured with your own script — it varies too much with script complexity to borrow numbers.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.