For example, having a folder of contacts with each file named after the person and having key/value pairs. Similar to how static site generators use YAML/TOML/JSON.
For example, having a folder of contacts with each file named after the person and having key/value pairs. Similar to how static site generators use YAML/TOML/JSON.
Or alternatively a midnight commander for databases. :)
When you need to store lots of files on disk, it's a common pattern to spread them out in subdirectories. For instance, instead of `files/2d8af74bcb29ad84`, you would have `files/2d/8a/f74bcb29ad84`.
That isn't how inode filesystems work -- if you change a file's permissions, it's just an inode update on the file -- not the containing directory.
Even in DOS-type filesystems (FAT/exFAT), it's just a record update in the corresponding dirent for that file.
If you add a new file to a directory, that causes an mtime update on the directory's inode.
The rest is accurate -- many older filesystems have lookup performance that scales poorly with directory size (for DOS filesystems and BSD UFS, you have to do a full directory scan). Also ls defaults to sorting output, which is O(N log N) and can be slow in large directories.
I'd also never put sqlite on NFS. Locking is often broken on NFS. Unless you can guarantee a heterogeneous environment. I can imagine some Excel guy in marketing is going to launch his SQLite UI on Windows and completely hose it all.
I think breaking down large collections in subdirectories is mainly done in order to ensure that it'll work even with FS that don't deal well with very large directories and also because it makes inspecting the files manually a little more convenient given that many file explorers (and especially the GUI ones) can have trouble with large directories.
I don't know exactly how Linux does it. Windows hands off whole paths to the filesystem, so this idea is possible there.
In FreeBSD, there is a generic routine (lookup(9)) that goes component by component, so at each step the filesystem is only asked to resolve a single component to a vnode. I think a clever filesystem implementation (in FreeBSD) could look at the remaining path and kick off asynchronous prefetch... but I am not aware of anything doing this.
There are two modes you can implement for your filesystem: in the so-called 'high level' mode you get the whole path. In the 'low level' mode the filesystem asks you for one piece of the path at a time.
The low level mode seemed faster in my tests, and I think it's also closer to how Linux kernel works internally?
I’d expect it to take up slightly more disk space. More importantly, the more indirections, the longer it’ll take to get to the file.
The problem is that most tools that operate on file systems doesn't. Things like readdir() is a linear scan and takes a long time on a million files.
So in practice it's not optimal. A thousand files, no problem. A million, start looking at doing it in-memory (or use some database tool).
Yes, depending on the file system.
For example, ext4 with default settings uses 32-bit hashes. Upon the first collision, you can no longer add more files to the directory (ENOSPC error).
Source: https://blog.merovius.de/2013/10/20/ext4-mysterious-no-space...
Perhaps they want to tag a few, and it’s useful to have some autodetected metadata but they don’t want to tag them all and they don’t want a gigantic ‘Untagged files’ list. They want folders and they want more than a flat folder list, they want nested folders.
You can implement it, it’s not hard and most modern file systems have all the features you need. But users will hate it and won’t use it the way you want.
But I am less sure users actually want folders.
Some power-users, sure. But most normal people don't want to deal with folders, either.
For evidence: look at the guy who saves everything on his overflowing desktop.
Any system that allows people to find their stuff, and perhaps make a few annotations, will be good for them.
Google Photos is almost a good example: I don't have to do annotate anything, yet I can search for eg pictures of snow or by location.
(I say only 'almost', because while impressive, that system isn't good enough yet to find obscure stuff or to work on contextual cues like 'those pictures I took at home after we came back from shopping sometime in the last few months'.)
And really, a lot of users don’t want an interface that doesn’t allow them to do what they want just because someone else just dumps all their files on the desktop.
Apple tried this on iCloud and had to go back. Because, while it makes for nice presentation and usability, there’s a lot of users that it can’t cater for.