Linux w/ reiserfs 3.x vs a SUN SAN 7410 running solaris and ZFS
badcheese.com
badcheese.com
Reiserfs does an "ls" in a directory with 6000 files in it in about 3-5 seconds. The SUN SAN does it in about 1-2 minutes. Serious problems here.
That has to be a bug, not merely slow filesystem behavior. Even doing a seek for every block (4k on ZFS?) at 10ms for a minute comes to 4kb of data per directory entry. That's ridiculous. Something is broken; probably an interaction between subsystems (hardware cache, software cache, ZFS filesystem, network filesystem, SAN configuration, etc...).
I'm not trying to be a troll, I'm trying to fix the performance problem.
mkdir /tmp/foo
cd /tmp/foo
for ((i=0;$i<6000;i++)); do touch $i; done
time ls >/dev/null
real 0m0.012s
user 0m0.012s
sys 0m0.000s
Perhaps try ext3?As for working with many small files, in my experience ZFS has been far better than xfs (I do not use reiserFS due to previous stability issues I experienced with it). One particular example was a user who had over a million files in one directory. This caused all our backup software to fail. After moving to ZFS, I could send and receive these files between servers without problem. I could list the files easily enough as well, after installing the aforementioned GNU versions with Blastwave.
I believe there is a GNU version of ls in Blastwave as well that Steve could try.
At any rate, it certainly doesn't take minutes to list 6000 files. My example above was actually on 4 million files. Extrapolating from 6000 to 4 million, my listing should have taken 11 hours. It may have with the default utilities, but it took no where near that long when I replaced the utilities with GNU versions.
For reference, this was directly on the Solaris machine which had the ZFS pool attached via SATA.
Not sure of a better general solution, but approaches like BigTable/GFS and SimpleDB come to mind.
Conversely if ZFS is supposed to scale to petabyte loads, what configuration do they expect that data to be in?
I guess instant, unlimited snapshots don't come free. But you also have the option of storing metadata cache on seperate storage (such as SSDs), a feature which many other filesystems don't offer.
real 0m 1.96s user 0m 1.12s sys 0m 0.00s
Sounds like Steve was having some other problem unrelated to ZFS.
EDIT: also just found a directory I had with 60,000 randomly created files over time (ie. fragmented), and ls took 3.5 seconds locally (didn't try it over NFS). This is looking more and more like a troll post :-)
Took me 9.9 seconds to get a directory listing for 65336 files over NFS after creating them over NFS on another system.
That's still no where near as bad as the author states, but I had those files in my cache on the file server, I bet.