When to Enable WebGPU Timestamp Queries for GPU Pass Measurement
Decision guide for enabling WebGPU timestamp queries: when it makes sense, how it compares to CPU timing and shader counters, and a safe implementation pattern with validation checks.
29 Mar 2026, 02:59 UTC

Decision Overview
Enable WebGPU timestamp queries for production profiling only when the adapter reports the timestamp-query feature and you need GPU-side timing of a specific pass rather than a wall-clock estimate. If the feature is absent, fall back to CPU timing around queue.submit or to shader-based instrumentation.
Constraints
- Adapter must expose
timestamp-queryinadapter.features. Availability varies by browser and OS and may be behind flags. - Each timestamp is 8 bytes. A resolve buffer must be sized for the number of queries and a separate staging buffer is required for MAP_READ.
- Readback requires a GPU to CPU copy and
mapAsync. Mapping stalls the pipeline until the data is ready. - Timestamp values are GPU timestamps in nanoseconds per the WebGPU spec. They measure execution on the GPU, not wall time, and are not comparable across adapters.
Option Comparison
| Option | What It Measures | Instrumentation Cost | Availability | Typical Overhead |
|---|---|---|---|---|
| Timestamp Queries | GPU timestamp at start and end of a pass | Low: one writeTimestamp per range, query set and resolve buffers |
Only with timestamp-query feature |
Small GPU overhead for resolve; CPU stall only on readback |
CPU Timing performance.now |
Wall-clock time around queue.submit |
None | Universal | Includes submission and scheduling latency; inaccurate for overlapping work |
| Instrumented Shader Counters | Custom counters written from shaders | Higher: extra instructions and register pressure | All adapters | Can change shader performance; needs resolve/readback |
Trade-offs
- Timestamp queries give GPU-side duration with minimal shader impact. They require feature support, a query set, a resolve buffer and a staging buffer. Frequent readback adds synchronization cost.
- CPU timing is always available and adds no GPU work, but it conflates command submission, queue scheduling and execution.
- Shader counters are portable but alter shader resource usage and may affect the measurement.
Implementation Pattern
Run in browser JavaScript after a WebGPU device is obtained. No special permissions are required.
// 1. Feature check
if (!adapter.features.has('timestamp-query')) {
// fallback path
return;
}
// 2. Query set for start and end
const querySet = device.createQuerySet({
type: 'timestamp',
count: 2
});
// 3. Resolve buffer and staging buffer
const resolveBuffer = device.createBuffer({
size: 2 * 8,
usage: GPUBufferUsage.QUERY_RESOLVE | GPUBufferUsage.COPY_SRC
});
const readBuffer = device.createBuffer({
size: 2 * 8,
usage: GPUBufferUsage.COPY_DST | GPUBufferUsage.MAP_READ
});
// 4. Encode work
const encoder = device.createCommandEncoder();
const pass = encoder.beginRenderPass({ /* ... */ });
pass.writeTimestamp(querySet, 0);
// ... render commands ...
pass.writeTimestamp(querySet, 1);
pass.end();
// 5. Resolve and copy for readback
encoder.resolveQuerySet(querySet, 0, 2, resolveBuffer, 0);
encoder.copyBufferToBuffer(resolveBuffer, 0, readBuffer, 0, 2 * 8);
device.queue.submit([encoder.finish()]);
// 6. Read back
await readBuffer.mapAsync(GPUMapMode.READ);
const data = new BigUint64Array(readBuffer.getMappedRange());
readBuffer.unmap();
const startNs = data[0];
const endNs = data[1];
const gpuDurationNs = endNs - startNs;
Placeholders /* ... */ indicate your own render pass descriptor and commands.
Validation Steps
- Confirm
adapter.features.has('timestamp-query')before creating the query set. Creation will be rejected without the feature. - Create a minimal pass with two timestamps, resolve and copy, then map. Verify the returned values are non-zero and monotonic.
- Increase workload size in a controlled way and check that measured GPU duration increases monotonically.
- Test on target browsers and adapter types to confirm feature presence and consistent behavior.
Limitations and Practical Checks
- Avoid mapping the read buffer every frame in production. Batch reads or sample periodically to limit stalls.
- Query sets consume GPU memory. Create only the number needed for the profiling window.
- Timestamp values are per-adapter GPU timestamps. Do not compare absolute values across devices, only relative changes within the same run.
- Feature support is implementation-dependent. Check the browser documentation for flag requirements.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.