Moving Heavy Loops to the GPU with WebGPU Compute Shaders
A practical walkthrough of WebGPU compute shaders: async pipeline creation, buffer and bind group layout, dispatch math, async readback, and the honest trade-offs versus plain JavaScript.
09 Jul 2026, 01:14 UTC

You have a loop in JavaScript that processes a million particles, pixels, or embeddings per frame, and the main thread is drowning. Web Workers help, but the work is fundamentally data-parallel: the same operation, millions of times, on independent inputs. That is exactly the shape of problem WebGPU compute shaders exist for, and getting one running is less exotic than it looks — provided you make a few engineering decisions correctly up front.
This post walks through those decisions: how to create the pipeline without janking the page, how to lay out buffers and bind groups, and how to get results back to the CPU without stalling. The example is a minimal element-wise computation, but the structure carries over to particle systems, prefix sums, and small ML kernels.
Check availability before you promise anything
WebGPU support is still uneven. Chrome and Edge ship it (Chrome 113+ on desktop), Firefox has it behind a flag or in progress depending on version, and Safari support has been landing through Technology Preview. So the first decision is a fallback strategy: WebGL2 compute isn't a thing, so your realistic fallbacks are WebGL transform feedback, WebAssembly with SIMD, or just a slower CPU path.
// Run in the page's main thread or a worker; no permissions needed.
if (!navigator.gpu) {
// Fall back to WASM/CPU path here.
throw new Error('WebGPU not available');
}
const adapter = await navigator.gpu.requestAdapter();
if (!adapter) throw new Error('No suitable GPU adapter');
const device = await adapter.requestDevice();
// Know your limits before allocating anything big.
console.log(adapter.limits.maxComputeWorkgroupSizeX); // often 256, sometimes higher
console.log(adapter.limits.maxStorageBufferBindingSize); // varies a lot by deviceDo not skip the limits check. maxStorageBufferBindingSize in particular can be far smaller than your data set on low-end devices, which forces you to chunk the work into multiple dispatches.
Create the pipeline asynchronously
Shader compilation is expensive, and device.createComputePipeline can block. The async variant, createComputePipelineAsync, lets the driver compile off the critical path — call it during page load, before the user hits the button that needs the result.
const module = device.createShaderModule({
code: `
@group(0) @binding(0) var<storage, read> input: array<f32>;
@group(0) @binding(1) var<storage, read_write> output: array<f32>;
@compute @workgroup_size(64)
fn main(@builtin(global_invocation_id) id: vec3<u32>) {
if (id.x < arrayLength(&input)) {
output[id.x] = input[id.x] * input[id.x]; // stand-in for real work
}
}`,
});
// Always handle rejection: WGSL compile errors surface here, not at createShaderModule.
const pipeline = await device.createComputePipelineAsync({
layout: 'auto',
compute: { module, entryPoint: 'main' },
});Two things bite people here. First, WGSL errors are reported asynchronously — wrap shader creation in device.pushErrorScope('validation') / popErrorScope() during development, or compile failures look like silent no-ops. Second, the arrayLength guard matters: dispatch counts are rounded up to whole workgroups, so the last workgroup will have out-of-range invocation IDs.
Buffers, bind groups, and the dispatch math
Data lives in GPUBuffers. For compute you typically need three: an input storage buffer, an output storage buffer, and a staging buffer for readback (storage buffers can't be mapped directly to the CPU unless created with MAP_READ, and you generally don't want your working buffers mappable).
const N = 1_000_000;
const byteSize = N * 4; // f32
const inputBuf = device.createBuffer({
size: byteSize,
usage: GPUBufferUsage.STORAGE | GPUBufferUsage.COPY_DST,
});
const outputBuf = device.createBuffer({
size: byteSize,
usage: GPUBufferUsage.STORAGE | GPUBufferUsage.COPY_SRC,
});
const stagingBuf = device.createBuffer({
size: byteSize,
usage: GPUBufferUsage.MAP_READ | GPUBufferUsage.COPY_DST,
});
device.queue.writeBuffer(inputBuf, 0, yourFloat32Array);
const bindGroup = device.createBindGroup({
layout: pipeline.getBindGroupLayout(0),
entries: [
{ binding: 0, resource: { buffer: inputBuf } },
{ binding: 1, resource: { buffer: outputBuf } },
],
});Dispatch size is the other classic off-by-one. With a workgroup size of 64, you need Math.ceil(N / 64) workgroups, not N:
const encoder = device.createCommandEncoder();
const pass = encoder.beginComputePass();
pass.setPipeline(pipeline);
pass.setBindGroup(0, bindGroup);
pass.dispatchWorkgroups(Math.ceil(N / 64));
pass.end();
// Copy results to the staging buffer in the same submission.
encoder.copyBufferToBuffer(outputBuf, 0, stagingBuf, 0, byteSize);
device.queue.submit([encoder.finish()]);
// Readback is async — batch your work so this latency is amortized.
await stagingBuf.mapAsync(GPUMapMode.READ);
const result = new Float32Array(stagingBuf.getMappedRange().slice(0));
stagingBuf.unmap();Run all of this in the page (or a dedicated worker) after device acquisition; no special permissions are required beyond a secure context (HTTPS or localhost). To verify it works, compare a small input — say 16 elements — against the same computation in plain JS. If the tail elements are wrong, your bounds check or dispatch rounding is off.
The trade-off you actually need to weigh
The uncomfortable truth: for one-shot work, WebGPU compute often loses to plain JavaScript. Buffer creation, pipeline compilation, and especially mapAsync readback each cost microseconds to milliseconds, and readback latency can dominate entirely. The win comes from amortization — either the data set is large enough that GPU throughput swamps the fixed costs, or the data stays on the GPU across many frames (a simulation feeding a render pass) so readback never happens at all.
If you only need the result on screen, keep it on the GPU: bind the output buffer directly into a render pipeline and skip the staging buffer entirely. That's where compute shaders really pay off.
Where to go from here
Start with the element-wise example above at a small N, verify correctness against a CPU loop, then scale up and profile. Chrome DevTools' GPU tooling and WebGPU timestamp queries (a device feature you must request explicitly) will tell you whether you're compute-bound or transfer-bound. Once the skeleton is solid, swapping the WGSL body for a particle integrator or a reduction is the easy part — the plumbing is the hard part, and now it's done.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.