Reliability of %util for detecting storage saturation in iostat
0 reputation · 29 Aug 2024, 01:48 UTC
Evaluating I/O Bottlenecks with sysstat
When analyzing storage performance using iostat -x, the %util metric indicates the percentage of time the device had at least one request outstanding. While this is a primary indicator for legacy spinning disks, modern storage architectures—including NVMe SSDs, RAID arrays, and virtualized cloud volumes—can process multiple I/O requests in parallel.
On these parallel-capable devices, a %util value of 100% does not necessarily imply that the device has reached its maximum throughput or IOPS capacity, as the device may still have available internal queues to handle additional requests.
To determine actual saturation, engineers often correlate %util with aqu-sz (average queue depth) and await (average latency). However, there is no universal threshold defined in the documentation to signal when a device is truly saturated versus merely active.
Which combination of await and aqu-sz provides a more reliable saturation signal than %util on parallelized storage? Does a rising aqu-sz coupled with stable await indicate available headroom despite high utilization?