Performing backups of our production apps used to take hours (especially in cheap clouds) because of all the loose files. Today, it takes about 3-5 minutes since there are just a handful of consolidated files to worry about.
Performing backups of our production apps used to take hours (especially in cheap clouds) because of all the loose files. Today, it takes about 3-5 minutes since there are just a handful of consolidated files to worry about.
And that might be better for accessing individual files outside the app instead of a big binary blob that you may have no clue on its format (or corruption/deletion of that single file). Thats why there are no universal solutions, there are many different use cases.
Nor does it let you start iterating through a directory with getdents(), then pause, then come back to iterate the rest of the files later. You have to receive every file name in the order the OS wants to give them to you.
That prevents applications really being able to make use of clever storage mechanisms - you can never find the most recent 3 files in a big directory with anything cheaper than a linear scan.
(But we have impressively speedy tools like Wiztree or Everything thanks to it.)
The most recent innovation for us is to use a lot of smaller SQLite databases, each scoped to a specific customer, session or unit of work. This is still far superior to loose records on disk, but you also get clear separation and reasonable firewalls between system entities.
There is zero reason a process would have more than tiny slowdowns with even millions of files in a folder. Finder has problems if you're trying to look at that folder, for obvious reasons, but it's a bit of a self-own for a backup co to claim that 200,000 files causes their solution to break. That speaks to serious algorithmic issues.
DISCLAIMER: This comment will be auto-dead because of moderation choices by dang (e.g. his pernicious need to pander to the anti-science, far-right crowd). This is a badge of honor. Never vouch for my comments.
I assume you are using Flex Tape as a derogatory comparison here. That said, I do view SQLite as a kind of "fix all" in the software world.
It's cheap (free), fast, everywhere and applicable to virtually every type of problem domain. AAA game assets to B2B line of business app storage. It's the most tested and used software on earth.
It doesn't fix everything, but it certainly gives you a fighting chance to make it to the next step.
Things like contact lists or game assets or even web history on individual computers will never grow to multi-TB sizes, hence no reason to over engineer them.
Might not be the most glorious or flashy solution, but like you mention, getting to the next step is all that really matters, IMO.
What kind of BS is this? Have you ever worked with one such folder? If you did, you would know that almost every app not doing some magic slows down to the point of being unsable (or even not working as described in original post). This is true for both Linux and Windows file systems.
a) 1M files in a single directory
b) 1k directories with 1k files in each
c) 1k aggregate files with 1k files in each
d) 1 aggregate file with 1M files in it
I think there's some filesystems that will work with option a, but any filesystem should work ok with the other options. Options c and d will make rsync much faster as you eliminate millions of syscalls.
It's usually a problem for any software that enumerates a directory for any reason, because that ends up being O(n) and often at least N system calls to e.g. get file information.