Show HN: Get lists of files in a directory that contains a large number of files
github.com
github.com
What makes "ls" appear slow on large directories is that it loads the entire directory to sort the contents. "-f" turns off that behaviour (it also changes some other switches; EDIT: another relevant change might be that ls by default also outputs columns that needs to know length of filenames), and "find" does not sort.
(Don't want to fault OP for writing something to do this - it's not obvious if you've suddenly had to deal with a huge directory and aren't quite familiar with trying find a way to list their contents faster - I remember my frustration over how slow that could be myself until I learned of those options, especially on ext2fs where it was really awful)
But I've run ls on systems with several orders of magnitude slower drives than what we're used to today, and it really isn't generally what causes people to complain of ls being slow on large directories in practice.
Even with ext2fs which was notoriously bad at handling large directories, exported via NFS filesystems over 10Mbs ethernet backed by 90's era IDE drives, the buffer size was rarely enough of a problem to matter once you turned off the sorting.
(To be clear, I'm not saying it won't ever make a difference; but in practice, over decades of running into this complaint, it's just not been the issue - most of the time peoples problem is interactive use where they run into having to wait for ls to read the whole dir to start getting results, not total throughput)
I'm sure there are circumstances where a 5x speedup will matter, but most of the time where people use "ls" it's interactively on the command line and it's just not been my experience that it matters in the circumstances where I've had to help people with this. It's more of a "we're right in the middle of something and can't get any results back" kind of situation. I'd say 9 out of 10 times I've had people bring up an issue like this it's because they're uncomfortable with "find" and are trying to get a list of files to do something with that they could easily do directly with "find" and/or where "ls -f" is more than sufficient.
But as you say, YMMV.
(also "-f" turns off the colors and type indication in any case, making it a more portable way of ensuring you get fast output even on systems which doesn't embed the type in getdents() results)
My solution was to use a nested directory approach where each could contain up to 4096 files or other directories. A path for the first 3 items in the sequence would look like:
/0/0/0/0.item
/0/0/0/1.item
/0/0/0/2.item
The 4096th & 4097th identities would live at: /0/0/0/4095.item
/0/0/1/0.item
Implementing the path scheme is a series of trivial bitmasks (0xFFF) over the identity of each.The only thing I don't like about this is the lock around ensuring the directory exists, but its not in an especially hot path. Updates to existing items are where things get tight for us, not on creation.
I typed a lot so i can be better corrected, not because I understand what is going on.
At the time people looked down on it mostly because everybody else did it differently and DOS looked like a toy OS compared to more advanced ones.
Of course, my comment is kind of a joke. There many advantages on more advanced ways to list files on a directory. Not having to make one syscall for every file is just one.
That is, glibc() seems aware that getdents64() is the right syscall in both of those contexts.
Either runs in 300ms or so for a directory with a million files on an old laptop via WSL2.
Is this solving a problem Linux used to have maybe? It doesn't seem to be a problem now so long as you call ls in a way that doesn't make it stat() every file. Or does using a larger than 32k buffer with getdents64() really make that much of a difference?
It's not solving a problem Linux used to have - I used "-f" or "find" to avoid this issue first time sometimes in the mid 90's.
But a lot of people aren't aware, and it's perhaps a poor default and/or it'd be nice if ls output a warning with a hint if a stat of the directory indicates it's likely to be a big one.
The buffer size probably does make a difference (and current glibc scales the buffer based on st_blksize reported from stat, up to 1MB at most), but on large directories "-f" and/or avoiding anything triggering a stat() will be a far bigger deal than upping the buffer size.
It would be nice if there was an easy way of adjusting the glibc buffer sizes though (currently it takes modifying constants in the source and recompiling glibc to increase it.
That's what I meant by "glibc seems aware that getdents64() is the right syscall in both of those contexts".
The likely culprit is stat - dentry doesn’t have all the attributes inline, so a stat call is needed for EVERY file. That’s a lot of syscalls. Even worse, in a large directory the information won’t be in a cache, so every single stat hits the disk, dominating the time spent.
You're right that the moment you give an option to ls that requires stat calls, though, the stat calls tends to dominate.
func blen(b [256]byte) int { for i := 0; i < len(b); i++ { if b[i] == 0 { return i } } return len(b) }
Is it not better to use a slice with 256 capacity and access to `len()` than for every file to iterate through the array until you find the end?
Btw: many tools choke on directories with a large number of files. E.g. Bash autocomplete and Vim autocomplete. And there typically seems to be no way to abort these operations with Escape or Ctrl-C or Ctrl-\. Perhaps these tools could benefit from the approach taken here.
You can list a directory containing 8 million files! But not with ls.
https://news.ycombinator.com/item?id=28191639 (128 comments)
EDIT: A reason to use getdents() over readdir() would be not having to stat to get the file type.
EDIT2: "ls" in GNU coreutils will actually get the file type from readdir() guarded by a feature flag to determine whether d_type is available on the platform in question, so clearly you can get d_type from readdir() too. No idea about accessing it from Go, though.