Why your ZFS pool is still fragmented even with recordsize set

Filesystem & Storage (ZFS/Btrfs/LVM)

Why your ZFS pool is still fragmented even with recordsize set

Technical Briefing | 9/9/2026

You carefully tuned your recordsize to match your database workload, thinking you had outsmarted the ARC. You check the pool, you see a decent hit rate, and you move on. But months later, your read IOPS are lagging behind the raw hardware specs and performance feels sluggish. The truth is that ZFS fragmentation isn’t just about disk layout; it’s about how the Adaptive Replacement Cache handles partially filled blocks when your write patterns don’t align perfectly with your storage pool topology.

The lie about free space

Most people look at zpool list and see 20 percent free space and assume they have headroom. That is a dangerous assumption. ZFS needs contiguous free space to write new blocks efficiently. If your pool is heavily fragmented, the allocator starts jumping all over the platters or NAND flash to find a hole big enough for a record. This turns your sequential writes into random seek fests, and no amount of RAM will fix that once the metadata cache starts thrashing.

zdb -mmd poolname

  • The fragmentation metric in zpool list is essentially useless for predicting latency spikes.
  • Look for the fragmentation value in the zdb output which shows the actual allocation state of your meta-slabs.
  • If you see high fragmentation, you might need to reconsider your pool layout or migrate data to clear out the mess.

I’ve seen this bite teams running high-volume log aggregators. They set a massive recordsize to save on metadata overhead, but the resulting fragmentation turned their SSDs into IO wait nightmares. Don’t chase the perfect recordsize if your application churns through small, non-aligned writes. Sometimes the default is the default for a reason, and if you must deviate, make sure your storage array is provisioned to handle the inevitable overhead of copy-on-write fragmentation. Keep an eye on those meta-slab stats before the performance degradation becomes your next production incident.

Linux Admin Automation  |  © www.ngelinux.com  |  9/9/2026

0 0 votes
Article Rating
Subscribe
Notify of
guest

0 Comments
Newest
Oldest Most Voted