Billion file filesystem
blog.liw.fi
blog.liw.fi
This can be useful up to a point. But be careful of over optimizing for the pathological case at the expense of typical usage. Most of the listed operations are usually done on a small number of files. And even when they seem a lot to the user it's still a teeny tiny fraction of one billion.
And that's before we even mention directories.
There are circumstances where a great many very small files are present: a mail archive (though with the amount of headers in a modern SMTP delivered message for host transit tracking and spam/other identification process notes, these files are not as small as they once were), similarly a usenet archive, a web cache, source repos, …. Also I've seen some tools use small files as per-user flags or session notes, where that don't want the extra dependency of a DB layer, to parse/update a more complex structure for on-disk updates, or the cost of serialising an in-memory structure in a large regular write. This results in very small or even zero-length files, and could balloon in numbers with a great many users. But it is getting more common for such things to be in a sqlite DB or similar instead.
--
[0] Ext4 can inline minute files, 60 bytes or less, in the inode if the relevant option is enabled¹ on filesystem creation (IIRC it can't be turned on after the fact), NTFS can inline small files (I think about 600 bytes) in the MTF structure.
[1] It isn't enabled by default because of a potential rare issue that results in the chance of corruption during recovery after an unclean unmount (due to power drop for instance). IIRC the conditions are: if you create a small file that is inlined, then append to it enough that it can no longer be inlined so a block is allocated, and the unclean unmount happens immediately after, the new data may be lost.
1. It takes so long to create that many files in the first place. In this case 26 hours of doing nothing but file creation.
2. If the files contained any data at all, you would need a lot of disk space. If each file was 64k, you'd need 640TB of disk. It seems more likely you'd run out of disk space before you reached a billion files.
3. Someone developing such a system where that many files could be created would probably realize that it's a poor design and figure out a way to mitigate it before it ever happened.
Nevertheless, it's fun to think about.
Not that many, no, but I can tell you from professional experience that real life Linux systems get really weird when you have merely tens of thousands of files in a directory, and it would be kind of nice if that were fixed.
I did find a tool that could do the job but of course no one wanted to pay for it. Fortunately for me I left that position before the migration and someone else had to figure it out.
So yeah, I’ve encountered something similar in the real world. It was the result of poor software design choices but by the time I got there that path was chosen and I had to deal with the results.
I can tell you, it took a real, real long time to delete seven million files on NTFS. It was not happy.
When it's not angry?
ls -U1 |wc -l
Can be very fast with lots of files on a modern Linux box. -U: do not sort; list entries in directory order find . -mindepth 1 -maxdepth 1 -printf x | wc -c
performs about the same and includes dotfiles (not "." or "..", though) in the count.Old Reiserfs3 really excelled in this, Gentoo Portage had many tiny files. Today Btrfs has features like extends, where tiny (or empty) files are stored together in single inode.
For distributing image with several files(such as this example), SquashFs is the best option.
Warms my heart
Put it this way - moving a sqlite file with 1 million rows is much easier than moving 1 million files.
Performance is "just" an engineering issue.
This is a tale and argument as old as time. It’s the squinting that is the problem. The devil is in the details.
There's also no ways to really make a file system faster when dealing with many small files and directories. If you can avoid doing it by using a database you probably should.
Depends on the file/directory structure. Traditionally back in the CGI days URI's were links to files on a server. They often still are for static files. If you know the path to a file then the lookup should be very fast.
But yeah any other type of index other then the path isn't really possible.
Sort of, yeah. Consider how many command line programs will accept a file as an input parameter vs. how many will accept an SQL query.
Existing tooling is predominantly designed to work with files so depending on what you're trying to accomplish, you may have a lot more work to do when using a database.
And yes - we work really hard to address the performance issues here, but there are constraints.
Or possibly a document based data store.
The monolithic persistence layer is an anti pattern unless you have specific reasons to support it.
ZFS is quite nice as filesystems go but it is still a filesystem.
Does anyone know if it is a custom engine/theme or not?
This kind of blog/website can easily be done by writing every blog post in a separate html file then doing a quick script to put together the index.
Your script could also add use a very simple template for each page, template in which you can use your own CSS. (here, it is an external file containing 142 lines of straightforward CSS).
Yes, I did it for my own blog. But I also added the fact that I would like to write my posts in the Gemtext format. So I wrote a parser to convert it to HTML.
And, in the end, it proved to be quicker and easier to do it all by hand myself than to learn any templating engine or any static website generator. Also, I never understood CSS (I’m really really bad at that) but managed to do what I want in exactly 42 lines of CSS. And, guess what, lot of people are asking me where I found my "theme".
I think that everybody working with the Web should be required to do a pure HTML/simple CSS website at least once to realize how easy and straightforward it is to NOT use any engine/theme.
And how much time and energy we are wasting on layers above layers above layers.
> Your script could also add use a very simple template for each page, template in which you can use your own CSS. (here, it is an external file containing 142 lines of straightforward CSS).
> Yes, I did it for my own blog [...]
So, you created a custom engine, which is what parent was asking about.
I could definitely be wrong, as it's still a bit ambiguous to me
Time comes out to 93.6 µs/file or 10 684 files/s created.
% b 26*3600*1000*1000/1000000000
93.60000000000000000000
Space comes out to 296 bytes per file. % b 276*1024*1024*1024/1000000000
296.35274342400000000000This isn't really a problem about optimizing the storage of a billion empty files, it's about what the file systems do with them.
However it still tickles my curiosity and exercises an interesting pathological case. As a kid I've had this idea, what if I could store the data bytes in the file names of empty files, did I just create space out of nowhere? It's always beautiful whenever you rediscover that computers are not magic.
The program output magnetization by timestep so for the simple structure I was investigating, I ended up with about 100K files of a few KB each which were a massive pain to move to my Linux machine for analysis.
I looked into changing the batching but couldn’t figure it out, and this prompted me for the first time to consider how the structure of data on a device could affect processing throughout even if the absolute quantity of data was the same.
no? it'd probably die soon.
> I’ve done this before, but this time I made it a little simpler for me to do it again: everything in one Rust program rather than clunky scripts.
The author then used this Python script: http://git.liw.fi/billion-files/tree/create-files
For a shell script version I would like to remind you that you can "touch" multiple files at the same time. There probably is a balance between the length of touch argument list (it takes time to generate it) and running more processes.
For readonly there are a couple good options, though arguably they are then archives. Many use sqlite because it's widely available and deals with this well. But most average file systems are pretty bad at at least one of those factors.
It's probably easier to identify which file systems are bad for which reasons. For example on NTFS you can't store a file in less than 1kb, and large directories can be an issue because a directory entry contains a sequential list of files. In contrast ext3 introduced a hash table to quickly look up file metadata in huge directories. On the other hand, if your files are e.g. about 800B and spread into directories of reasonable file count NTFS might suddenly outperform ext4 on some metrics, because NTFS can store the first ~900 bytes of the file together with the metadata, ext4 only the first 60. Or if your huge number of tiny files contains lots of identical files a zfs with live deduplication might become interesting.
Ah, but it has its own consequences. ext4 doesn't dynamically reclaim inodes or garbage collect these directory hash tables, so if you create a directory and cycle a bunch of files through it, like millions and billions of files many times over, you can get into a case where you have sizable portions of your drive as free space, but you get ENOSPC errors because adding files to the directory tries to insert them into the hash table, which fails, because the hash table takes up so much space it causes the drive to be actually full. Oh, don't try actually doing directory operations without the index, either. Something like `du -sh` without that index will take upwards of 10+ minutes on a fast virtio SSD to account for like 10 gigabytes of files.
I recently had a 1TB ext4 drive that got into this state where it had 150GB free but the directory hash indexes were fucking massive, because billions of files had gone through the filesystem and it was basically hosed because even turning off the index and then turning it back on does not drop the hash table index, it only literally turns off the index use in the read path and leaves the index as is. I couldn't find a way to drop and clear the hash indicies. Yeah. Apparently you can get into a similar state if you exhaust the inode count because, again, there is no dynamic inode reclamation.
In this case, I had the `gecko-dev` Git repository, which is absolutely massive, and I was extracting many copies of it onto my drive as part of some testing automation. So that added up fast. It was on a separate drive from my main /home mount, at least.
The lesson I have learned from this is to mostly just use XFS instead of ext4, I guess. Which I was already doing on all my servers, so a bit funny I learned that lesson only on my own personal machine.
So definitely not read-only, not identical, not all tiny 0b or even 800b files (maybe <1mb is still tiny?) , and storage overead is a much lower priority vs latency of access and speed of operations
Your question reminds me of a famous (some fifteen-twenty years ago) question from MSDN forums.
A poor soul complained that when they put 10K buttons on a form in the designer view in MSVS it comes to a grinding halt. And that's impossible to use the studio anymore. To which a bunch of respondents quickly replied: "but why tho? why do you need so many buttons on the form?" -- From the discussion that followed it became apparent that the poster wanted to reproduce the minesweeper game with a 100x100 board.
What I'm trying to say is: if you have a problem where you think you need a billion of files all in one directory, it's easier to reformulate the problem than to try to find a solution for it.
Now we are working on scooping up a bunch of small (1-2KB) files we want to index with metadata and chucking them into a SQLite file, because as fast as `grep` is searching through the text the constant open() and close() calls are much more impactful.
Systems I've built generally use a pointer within a database to a secondary location. I've also seen hundreds of megs stored in binary json blobs.
More recently our map tiles are limited to a single US state, and we are using XFS, but because of the much smaller data set size I haven't really done any optimization beyond hard linking duplicate files. I inherited that choice, so I don't know what went into that.
I wonder why one needs to use as much as 296 bytes per empty file.
File systems like ext4/NTFS/APFS/etc require extra disk space for "housekeeping" of data blocks and allocations and metadata (lastwritetime, permission flags, etc) . This reserved space for each file has a minimum fixed amount even if the file's size is zero.
E.g. ">By default, ext4 inode records are 256 bytes," -- from https://ext4.wiki.kernel.org/index.php/Ext4_Disk_Layout#:~:t...
(Scroll up on that page to see all the fields of metadata in each inode record.)
At the cost of introducing a Y2038 issue. The timestamp fields in the first 128 bytes of an ext2/ext3/ext4 inode are only 32 bits each, they are extended with an extra 32 bits each in the next 128 bytes of the inode.
Seaweed uses 40 bytes of metadata per file.
good ol easy-mode full-day task
I think that if some of the setup (mkfs, mount) is left to external tools¹, it show clearly how simple a program in Rust can be.
Sure, in Ruby, python, or even bash, it would be two or three lines, under ten statements and probably under 200 characters while remaining readable. But still it shows us how the idea that "in rust there is a lot of boilerplate" isn't really a practical issue.
I write most of my simple tools in rust lately. Needed a million URLs parsed and matched against denylists last week: my Python code took 5 minutes to run that (way too long for trial-error), the bash cut,sed,awk,grep pipe-chain quickly too wieldy to keep playing around with (yet seconds to run, not minutes). In a few minutes I was running a rust version that used all my cores, well readable and easy to iterate with. I guess go would be the perfect fit, but I don't speak go that well (nor have its tooling ready and up-front).
¹something I'd do regardless. If only to follow the part of the Unix philosophy "do one thing".
Something like, in bash:
test -f output.img || (fallocate -s 1T output.img && mkfs.ext4 output.img)
mkdir -p toto
mount output.img toto
for i in {1..1000000000}
do
touch toto/$i
done
umount toto
Or am I missing something ?http://be-n.com/spw/you-can-list-a-million-files-in-a-direct...
% time ls | wc
10000000 10000000 78888897
ls 1.90s user 0.73s system 99% cpu 2.637 total
wc 0.28s user 0.02s system 11% cpu 2.597 total
ETA: and most of that time is spent sorting the output. % time ls -U | wc
10000000 10000000 78888897
ls -U 0.53s user 0.29s system 99% cpu 0.823 total
wc 0.26s user 0.02s system 33% cpu 0.822 totaltime (real user sys) in seconds with 1e6 files in a xfs filesystem. Directory used 27 MB space.
ls -l > /dev/null : 49.12 23.90 13.47
ls -lU > /dev/null : 46.62 33.11 8.01
ls > /dev/null : 6.36 1.97 3.63
ls -1 > /dev/null : 2.49 1.83 0.36
ls -1U > /dev/null : 0.31 0.18 0.08
(find) : 0.77 0.49 0.19
and for 1e6 files in an ext4 filesystem. Directory used 23 MB space.
(but as i had no ext4 fs with one million free inodes on a loop device) ls -l > /dev/null : 51.39 34.02 11.39
ls -lU > /dev/null : 47.39 30.72 11.09
ls > /dev/null : 4.83 3.56 0.65
ls -1 > /dev/null : 4.46 3.45 0.50
ls -1U > /dev/null : 0.51 0.18 0.26
(find) : 1.29 0.76 0.36Of course I was not able to reproduce this difference and assume that I did work on the system and it was not idle.
I have repeated the the tests, on another mostly idle system. (Therefor the timing is not comparable to the first test). This time both filesystems live in a loop device and the benchmarks were run with hyperfine (3 warmup loops, 3 measurements each).
XFS
ls -l >/dev/null Time (mean ± σ): 5.890 s ± 0.394 s [User: 1.096 s, System: 3.491 s]
ls -lU >/dev/null Time (mean ± σ): 2.254 s ± 0.386 s [User: 0.648 s, System: 1.602 s]
ls >/dev/null Time (mean ± σ): 542.7 ms ± 4.4 ms [User: 443.8 ms, System: 98.7 ms]
ls -1 >/dev/null Time (mean ± σ): 534.4 ms ± 2.3 ms [User: 450.1 ms, System: 84.1 ms]
ls -1U >/dev/null Time (mean ± σ): 115.8 ms ± 1.8 ms [User: 80.9 ms, System: 34.6 ms]
ls -U >/dev/null Time (mean ± σ): 121.0 ms ± 14.6 ms [User: 80.9 ms, System: 40.0 ms]
find >/dev/null Time (mean ± σ): 225.0 ms ± 16.2 ms [User: 150.1 ms, System: 74.7 ms]
EXT4
ls -l >/dev/null Time (mean ± σ): 10.493 s ± 0.392 s [User: 2.495 s, System: 5.032 s]
ls -lU >/dev/null Time (mean ± σ): 8.395 s ± 0.157 s [User: 0.735 s, System: 5.049 s]
ls >/dev/null Time (mean ± σ): 1.966 s ± 0.010 s [User: 1.760 s, System: 0.202 s]
ls -1 >/dev/null Time (mean ± σ): 2.014 s ± 0.032 s [User: 1.798 s, System: 0.216 s]
ls -1U >/dev/null Time (mean ± σ): 221.4 ms ± 1.3 ms [User: 83.7 ms, System: 137.6 ms]
ls -U >/dev/null Time (mean ± σ): 222.5 ms ± 2.0 ms [User: 78.9 ms, System: 143.6 ms]
find >/dev/null Time (mean ± σ): 531.0 ms ± 4.0 ms [User: 377.3 ms, System: 153.7 ms]However, that was not the point of my reply
Though, I understand the author wasn't claiming to deliver an example of how to do "scripting" in Rust.