Note that Linux ext3/ext4 have a dir_index option which improves readdir performance.
Edit: it's not readdir performance so much as just name lookups which are improved by this technique (and dir_index): http://ext2.sourceforge.net/2005-ols/paper-html/node3.html
- I never needed to scan the directory to find "all users" (after all, if you have millions of users, this is going to take a while whatever directory structure you use)
- Modern filesystems either use a tree or hash structure to identify files, meaning that lookup by name, and creating/deleting files, is quick, even if you have millions of files.
- Given no performance benefits are to be had by directory nesting, I always went with the option of simplicity, i.e. having everything in one directory.
(I blogged about this here: http://www.databasesandlife.com/flat-directories/)
But no doubt the HN developers had a reason for doing this change, I'd love to know what it is (e.g. if they need to do something I never needed to do, or if they need to do the same things but I was wrong.)
But traditionally, ufs and ext2/3/4 (without dir_index) have to perform a linear scan through a linked list for lookup, and so they do indeed grow slower with number of files. This is likely where the fanout strategy originated from.
So as usual, YMMV and you should test on your file system of choice.
Personally, I don't really consider that fanout adds much complexity and I'd be surprised if it hurt performance.
edit: HN runs on FreeBSD. Not sure if they are using zfs or ufs, but I'm going to guess ufs. UFS apparently has a dirhash which improves directory lookups, but it's an in-memory structure so it won't help in the cold-cache case after reboot and it can be purged in low memory situations too.
https://wiki.freebsd.org/DirhashDynamicMemory
edit2: I wonder whether the HN admins ever tried tuning the dirhash settings? http://lists.freebsd.org/pipermail/freebsd-stable/2013-Augus...
> I never needed to scan the directory to find "all users"
> lookup by name, and creating/deleting files, is quick, even if you have millions of files.
And a question: given these observations, where do the benefits of filesystem fanout come from? Is it not true that looking up a file by name is fast no matter how many other files sit in the same directory? Is HN doing something weird?
You can't answer the question "where do the performance benefits come from?" by saying "look, the performance benefits exist".
I think he is trying to say is that the parent poster's observations must be wrong. After all, we are talking about an unsubstantiated claim ("there's no benefit to fanning out files") that directly contradicts another claim which we have data for ("HN is 5x faster after fanning out files").
The comment I was replying to was saying that the file system takes care of it automatically, so there's no purpose to arranging millions of files into directories. I'm not going to speculate how it all works under the hood.
I was once i charge of a large number of images of book jackets named by the books' isbn. At least in that population (which of course is an extreme example considering how an isbn is created) the distribution is much more even (that is, the directory sizes are relatively equal) when using the end than the beginning, but I would not be surprised if that is a normal outcome.
Maybe it's an application of Benford's law [1].
Some filesystems are worse at this than others (xfs... let's not go there).
(And if it did need to, presumably it's now need to recursively scan all sub-directories, which would also take a while?)
The same way a database fetch doesn't load the whole table, filesystems can and do use trees and hashes to organize directories so that file lookup, creation and deletion by name can be fast and can be concurrent.
I posted this in another comment, but this was my understanding of the situation in 2010. http://www.databasesandlife.com/flat-directories/
There must be a reason why they did this change, either I am wrong about performance (perhaps my results really were particular to those filesystems) or perhaps I am right and they made the change for another reason. I'd like to learn the answer.
For one it would need a vast amount of temp space, the site would be down while doing it and the end result would be much the same as what it is today (I rarely modify the filesystem).
EDIT: bzbarsky's explanation below is more accurate.
Instead of: command *
for i in [someregex]*
do
command $i
done
I know I could also do command [someregex]* but like the comfort of having each item echo back to the terminal so I know the progress.
So a million files in a dir is going to take longer to access any individual file than if there's only 3 files to pick from.
And if it scales worse than linear, a tree structure, although hitting the FS multiple times once at each level, in total can take less time.
Finally if you can avoid a smooth distribution hash and intentionally order by something important (time?) then you only need a cache the most recent directories in memory and the deep historical archive can fend for itself rarely accessed without getting in the way of the busy files. If you rarely if ever leave /stuff/thisYear/today/ then whatever is in /stuff/2011/dec25 will never slow today down or get in the way.