Stop your eBPF metrics from lying to you about interrupt latency
By Saket Jain Published Linux/Unix
Stop your eBPF metrics from lying to you about interrupt latency
Technical Briefing | 10/7/2026
We have all been there. You are looking at a system with high CPU iowait, but the disk latency metrics show everything is fine. You check the usual suspects and find nothing. Eventually, you realize the kernel’s scheduler is getting starved because of unmonitored softirq storms. The default tools just group this as part of system time, making it look like your application is just being slow when it is actually being pushed off the CPU by hardware interrupts.
Why top and htop get this wrong
The standard procfs-based tools track process states by scanning the task list. But interrupts happen outside the scope of your typical task context. If you have a NIC firing thousands of interrupts per second, you are spending cycles just context switching into the hard irq handler. Most monitoring agents ignore the IRQ count entirely, and that is how you end up chasing phantom performance issues for three days while the actual bottleneck is a poorly distributed interrupt affinity.
watch -n1 'cat /proc/interrupts | awk "{print \$1, \$NF}" | sort -n'
- Check which CPU core is handling the bulk of the traffic by monitoring the columns in /proc/interrupts.
- Look for a single core running at 100 percent while others are idle as a sign of unspread interrupt load.
- Consider pinning interrupts to specific cores if your network throughput is causing jitter for your database threads.
If you see one CPU core doing all the heavy lifting, you need to check your smp_affinity settings. It is often a side effect of leaving the default interrupt coalescing enabled on hardware that does not handle high burst traffic gracefully. Do not just blindly turn on irqbalance and walk away. Check the stats, verify the core distribution, and keep an eye on the context switch rate in vmstat before you assume your code is the problem.
