Diagnosing and Resolving High CPU I/O Wait (%iowait) in Linux
Learn how to diagnose high %iowait in Linux. This guide covers identifying blocked processes in the 'D' state, analyzing disk saturation with iostat, and resolving bottlenecks caused by atime, swapping, or hardware failure.
23 Oct 2025, 05:19 UTC

The Problem: CPU Idle but System Unresponsive
You observe high CPU utilization in top or htop, but the actual user-space processing (%us) and system-space overhead (%sy) are low. Instead, the %wa (iowait) value is high. This indicates that the CPU is technically idle, but it cannot proceed because it is waiting for outstanding disk I/O requests to complete.
The primary takeaway: High iowait is not a CPU problem; it is a storage bottleneck. The CPU is simply reporting that it has nothing to do until the disk responds.
Quick Diagnostic Matrix
| Symptom | Likely Cause | Key Metric to Check |
|---|---|---|
| Constant high %wa, low disk throughput | Slow disk/Network latency (NFS/EBS) | await (iostat) |
| Spiky %wa, high disk throughput | Application I/O saturation | %util (iostat) |
| %wa spikes during memory pressure | Excessive swapping (Thrashing) | si/so (vmstat) |
| Intermittent %wa with system hangs | Hardware failure/Bad sectors | dmesg logs |
Step-by-Step Diagnostic Workflow
1. Identify Blocked Processes
Check for processes in the D state (Uninterruptible Sleep). A process in this state is usually waiting for I/O and cannot be killed even with kill -9 until the I/O request returns.
# Run as root or user with process visibility
ps aux | awk '$8 == "D" { print $0 }'
Expected Result: A list of processes currently blocked on disk operations. If the list is long, you have a systemic I/O bottleneck.
2. Analyze Disk Saturation and Latency
Use iostat (from the sysstat package) to determine if the disk is physically saturated or if the requests are simply taking too long to complete.
# Run on the host; 1 second interval, extended statistics
iostat -xz 1
Focus on these columns:
- %util: If this is near 100%, the disk is saturated. The device cannot handle more requests.
- await: The average time (in milliseconds) for I/O requests to be served. If this is significantly higher than your hardware's rated latency (e.g., >10ms for SSDs), you have a bottleneck.
- avgqu-sz: The average queue length. A high number indicates requests are piling up.
3. Check for Memory Thrashing
If the system is out of RAM, the kernel moves pages to the swap partition. Because disk is orders of magnitude slower than RAM, this causes massive iowait.
# Run on the host; 1 second interval
vmstat 1
Look at the si (swap-in) and so (swap-out) columns. If these values are consistently non-zero during iowait spikes, your problem is memory exhaustion, not necessarily a failing disk.
Fixes Based on Findings
Scenario A: High %util and High await (Disk Saturation)
If the disk is saturated by a specific application, consider reducing unnecessary write operations. A common culprit is the atime (access time) update, where Linux writes to the disk every time a file is read.
Fix: Remount the filesystem with noatime.
# Run as root
# Temporary remount to test impact
mount -o remount,noatime /var/lib/mysql
Risk: Some legacy mail servers or backup tools rely on access times. Verify application compatibility before making this permanent in /etc/fstab.
Scenario B: Low %util but High await (Latency/Hardware)
If the disk isn't "busy" but requests are taking forever, check for hardware errors or network timeouts (if using NFS/iSCSI).
# Run as root
dmesg | grep -iE "error|timeout|scsi|ata"
Fix: If you see "I/O error" or "resetting adapter," the hardware is failing. Replace the drive or check the cable/fabric.
Scenario C: High si/so (Swap Thrashing)
Fix: Increase available RAM or tune the swappiness parameter to make the kernel less aggressive about swapping.
# Run as root
# Reduce swappiness (range 0-100; lower means less swapping)
sysctl -w vm.swappiness=10
Verification and Escalation
To verify the fix, monitor the %wa column in top and the await column in iostat -xz 1. The iowait should drop, and the processes in the D state should disappear.
Escalate to Hardware/Infrastructure teams if:
dmesgshows persistent sector errors.awaitremains high on a cloud volume (EBS/Azure Disk) despite low%util, suggesting you have hit the IOPS limit of the provisioned volume tier.- The bottleneck persists after disabling
atimeand resolving swap thrashing.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.