Stop your disk I/O wait from hiding behind average latency

Performance Monitoring & Tuning

Stop your disk I/O wait from hiding behind average latency

Technical Briefing | 7/29/2026

You see high iowait in top or htop and instinctively reach for iostat. It tells you your disk is doing 500 IOPS at 10ms latency, which looks perfectly healthy. But your application is still hanging on file operations. This bit me in prod during a database migration, and it took me two hours to realize that I was looking at the mean, which completely masked the outliers that were actually choking the queue.

The mean is a lie

Averages are great for making graphs look smooth in Grafana, but they are dangerous for debugging performance. Disk latency distributions are rarely bell curves. They are usually heavily skewed, meaning your 10ms average is actually a mix of a thousand 1ms requests and one 2-second stall. That single long request is what your application sees, and iostat just glosses over it entirely.

biolatency -D 5

  • biolatency from the bcc-tools package shows you the actual histogram of disk latency
  • Watch for the long tail where your requests are hitting the 100ms or 1s buckets
  • Check if these spikes correlate with specific processes or filesystem operations like metadata journaling

If you see a spread where most requests are fast but some are taking seconds, you aren’t looking at a throughput problem. You are looking at a contention problem, often caused by noisy neighbors or poor scheduling. Don’t trust your dashboard averages when the logs show clear delays. Fire up the histogram tools and actually see what the hardware is doing, because the kernel’s default reporting is just trying to be polite.

Linux Admin Automation  |  © www.ngelinux.com  |  7/29/2026

0 0 votes
Article Rating
Subscribe
Notify of
guest

0 Comments
Newest
Oldest Most Voted