Faster than the filesystem (2021)
sqlite.org
sqlite.org
That said, if you're using a file system abstraction for complex and compound documents using a database is a really stellar way to go. In part because it doesn't have the "chunking" problem where allocation blocks are used for both data and metadata in file systems so you pick a size that is least bad for both, versus one "optimum" size for disk io's to keep the disk channel bandwidth highly utilized and the naming/chunking part as records inside that.
I wrote a YAML <=> SQLite tool so that apps could use an efficient API to get at component types but the file could be "exported" to pure text. And it worked well in non-UTF string applications. (this was fine for the app which was orchestrating back end processes) At some point it would be interesting to move it to UTF-8 to see how that worked out.
I think that we stuck in 80-s/90-s foundational tech for too long and there're many new ideas that will shape future computing.
The endgame is a machine with only an OS. The OS is where all the hard edge cases are, which prevented the db-as-fs idea from succeeding in the first place.
What does this mean?
But of course, even in the context of the cloud a DB does not (necessarily?) replace a FS, and, more fundamentally, that which the poster presented looks more like a sociologistic presentation of trends (or, "big ideas" for the salesforce), oblivious of the absiological ground that "progress is not in decreased freedom", while the presented model is that of crippled machine (wanting to store locally is a basicmost demand).
This was tried multiple times already, and in each case, it turned out that the ability to easily store and organise data locally was both desired and required among the wider audience. Cloud storage is a nice feature to complement local storage, but will not replace it any time soon.
I’m sure there is some consumer technology coming around the bend that will make local storage again compelling but I think we are on the cusp of another server centric file storage era not the reverse.n
I can easily see a db file system being the default in for remote resources in consumer devices while traditional local file systems are reserved for specialized use cases.
It's faster, cheaper and often more reliable than someone else's computer.
And because of privacy concerns. And because people want to use their computers even when they cannot connect to the internet. And because people want the security of being sure they can access their files even if a company goes bankrupt/has an outage.
And because, no matter how fast networks get; I somewhat doubt that it will be faster to load a 2GiB 4k video file or a 7GiB pytorch checkpoint from cloud storage, than via the SATA or PCIe bus.
That said I'm now in this weird mid-life confusion about the fact that my new WAN (3gbps fiber connection) is now faster than any reasonable home LAN I can set up. Needed to move some movies from one machine to another the other day and it was way faster to just bittorrent them again from upstream than actually copy, since even USB removable media was slower than my WAN. So I think we might be at least temporarily in an awkward situation where cloud storage outcompetes LAN storage.
But truly local storage (nVME etc) has an extreme edge over anything networked, and always will.
Sure, for specific usecases that don't much resemble what most people spend their time doing.
That still doesn't make it faster to just torrent them again.
If the LAN is slower than the WAN then a large file where the transfer time is long enough that the few milliseconds of setup latency is irrelevant might be effectively equally fast to re-download from than to transfer over the LAN, but never faster, and in most cases the LAN is going to be more reliable.
There are almost certainly some cases where BitTorrent can transfer a specific set of files faster than common LAN transfer protocols, especially FTP or older versions of SMB when faced with large collections of small files, but that's an entirely different matter.
If the bottleneck in either case is the second PC's LAN interface how could the WAN possibly be a faster source than the LAN machine?
Granted, it might just be sub-optimal in certain scenarios, but I already have to logically partition my data into big blobs that are rarely accessed and stuff that must go into the SSDs already. Maybe just giving that FS a hatch to deal with the stuff that must be special and odd/weird is ok. I don't really want the DB magic for my swap files anyway.
The BeOS file system, an OS geek retrospective | Ars Technica
[https://arstechnica.com/information-technology/2018/07/the-b...]
http://www.nobius.org/dbg/practical-file-system-design.pdf
(Legal source, the website is the author's.)
Dominic Giampaolo now works at Apple on APFS.
He lives in Maine, though, so he did it over SSH!
If a file is short, the data doesn't need a whole disk block allocated to it. Likewise, if a file name is really long, that string doesn't need to be stored together with other metadata.
I don't want to remember paths and folders and all that crap. I don't want to depend on my discipline to organize and I'm willing to trade some flexibility in return.
There's some alternatives like tags, but none of them feel natural and require setup. I feel services like dropbox had a real chance to try something, but the best they came up with was showing me "recent documents" for a homepage, wasted opportunity if you ask me.
This is made even more difficult by the practical necessity of backward compatibility with standards like POSIX, SQL, et al in the same implementation that were never intended to be interchangeable at the implementation level.
In principle you could design a database to be used as an effective filesystem. In practice, the implementation wouldn't be very compatible with any other software you need to work with and the ability to integrate matters.
But I also have traveled enough roads in this industry to know that heterogeneity of approaches ends up defining actual practice -- for better and for worse -- and the "universality" of fopen/fread/fwrite/fclose of blobs of data is going to be hard to move away from.
I'd say that the only way that the "database" takes over the filesystem is by basically dropping to the lowest common denominator and becoming more like a filesystem; basically a key-value store. Throwing away the benefits of what a good database offers. Mainstream databases (in the form of SQL) only do a subset of what's possible with the relational model that Date&Codd laid out decades ago. I fear a move into a lower level of the stack would only worsen this.
In fact I think the trend is kind of going the other way. Filesystems are slowly adopting more and more of the storage techniques that came out of DB research, but using them for the fairly anemic storage Unix FS storage model. Because that's basically what the industry is asking for.
Further, I think there remains a poverty of knowledge about databases generally in the industry. I don't have a CS degree myself, but almost everyone I know (outside of my current employer) who does have one has said some variant of this to me: "I didn't take any/many DB classes in school." Knowledge of the relational data model is weak/thin in our industry (the low point being the "NoSQL" wave a decade ago). And within CS academia, it seems that R&D in this field is mostly centred around Europe (with the notable exception in North America being the excellent programme @ CMU.)
Disclaimer: Employed by a database company.
https://www.youtube.com/watch?v=dWIo6sia_hw (file system functions) https://www.youtube.com/watch?v=Va5ZqfwQXWI (benchmark vs SQLite) https://www.youtube.com/watch?v=OVICKCkWMZE (benchmark vs Postgres)
The underlying datastructures are really the same in a filesystem and a database in every way that matters.
Wouldn't it have been possible to not use a one-size-fits all approach? I'm thinking about things like moving the swap to a dedicated partition, like on Linux, and ditto for the other edge cases.
Nothing is stopping you from storing blobs of data inside a database and then exporting a POSIX api and calling that a new and improved filesystem. But once again its hard to see what value you get from all this complexity. A filesystem is complex enough and it doesn't have to store structured data. You generally don't want the OS to handle this complexity; you have just added another failure mode to a part of the system that you really don't want to fail.
1. Make it a database with schemas. In that case, you'll never be able to change or extend the language used to define schemas, since every application and OS will need to keep supporting older versions. It will also be hard to have multiple competing implementations of a complete database system, particularly because applications will come to rely on the performance characteristics of the reference implementation.
2. Make it schemaless, like a key-value store. That basically just is the file system, but non-heirarchical and with a faster implementation. In that case, your DB could just be a faster implementation of the existing filesystem interface or a natural extension thereof. The hierarchical nature of the filesystem is fairly essential if you want to allow multiple applications to avoid trampling on each other.
A file system is, at its heart, a naming system. Given this name, return me a handle to "do things" with that object. In SQL it is "SELECT * FROM FILES WHERE ( NAME == <name> );". The next debate is what the schema for the files table?
Well known things that would appear in that schema are things like "access rights", "ownership", "access times", Etc. Additionally related things like the "consuming application" might be there, and the "editing application" Etc.
If these things were all relations/schema then a lot of "warts" like MAGICNUMBER or extension type, security certificates, integrity digests, OS requirement, symbol tables, or character encoding can then become schema elements and the names get "purified" (scare quotes because architecture astronauts really get all hot and bothered by overloading things with multiple semantics.
Then you can do "views" like when someone types a name at the shell you can select from files that are executable on this OS and match the names.
There are basically a whole bunch of things that are "grafted on" to the file system and they can (and do) get out of date relative to the files and cause problems. (the canonical cases are moving an executable from one directory to another making it no longer visible for execution, changing the extension or magic number resulting in the wrong thing being used to try to run/edit it)
The architecture reasoning goes, "If there was just one source of truth about these things, a database, then a whole bunch of bugs and user annoyances would vanish."
Anyway, I've witnessed people start down this grand vision, devolve into your #2 above (key/value store) and then finally throw their hands up in surrender.
One of the systems questions I ponder sometimes is at what point is memory so cheap and plentiful that replacing the buffer cache in the kernel with a giant interconnected schema like this gives equivalent "time to record" values. An RDMA accessible victim cache for "blob" buffering might help too.
A couple of people have taken runs at it as "object storage" or "storage as a service" but it isn't quite there yet.
Still, for a complex file format (say MPEG 4) having it be a data base gives you some advantages and makes writing file component parsers unnecessary. So that's a win.
Edit:
Another thing that is sometimes forgotten when comparing SQLite to the filesystem is that files are hard[1]. It's not only about performance, but also about all the guarantees that you get "for free", and all the complexity you can remove from your codebase if you need ACID interactions with your files.
[0] https://en.wikipedia.org/wiki/Read%E2%80%93modify%E2%80%93wr...
Note that some filesystems have mount options that let you remove some of those guarantees.
IIRC the problem was its performance. I don't have any insider knowledge so cannot pinpoint the culprit but I suppose that the performance issue was probably not something fundamental tradeoff (as this article suggests) but more of its immature implementation. The storage technologies got much better nowadays so many of its problem could be tackled differently. Of course the question it has to answer is also different; is it a still worth problem to solve?
Whatever overhead of such small files might not matter though if the problem space is kept “human sized” (whatever one human can be bothered to manage).
I used to have tens of thousands of file in git annex. I had to tarball (chunk) some of the things that I never really use in order to speed up `git annex fsck`.
However, things have totally changed with the M1 generation of Macbooks. Things that were once near impossible now run in an instant. I need to redo this experiment.
IIRC, the fix was for directory entries were later stored in sorted buckets (aka hash map) and only the corresponding bucket for the file name would get locked. This reduced scope of lock for atomic operations like rename and also allowed faster lookup instead of O(n) based scan.
One core feature is that Windows offers hooks ('filters') that allows other components to put themselves between the client program and the files. This is how virtual filesystems work (like OneDrive, etc.), or how anti-malware works.
When you read from sqlite, then those reads can't come from another server, the objects can't be scanned automatically, etc. Again -- you might not need these features, but it's not that sqlite or ext4 is somehow magically faster; they just made different design choices.
That's not to say filesystems couldn't be improved.
Also be careful if the function that returns directory contents sorts results that you don't need done.
> SQLite reads and writes small blobs (for example, thumbnail images) 35% faster¹ than the same blobs can be read from or written to individual files on disk using fread() or fwrite().
> The performance difference arises (we believe) because when working from an SQLite database, the open() and close() system calls are invoked only once, whereas open() and close() are invoked once for each blob when using blobs stored in individual files. It appears that the overhead of calling open() and close() is greater than the overhead of using the database.
This should come as a surprise to no one. The rest of the article is only of interest in terms of how benchmarking is done, if that's your thing.
Off the top of my head, the typical filesystem stores:
- content
- creation time
- modification time
- last access time
- read/write/execute permissions
- owner
- group
- position in the dir hirarchyLazytime is a pretty good hack, but it doesn't survive a crash, does it?
https://github.com/guardianproject/libsqlfs
or something similar to check the potential for speed up.
Edit:
Holy shit, IO rings released on Windows Preview in 2021...
And yeah, io_uring helps since open-write-close can now be condensed into a single syscall.
The overhead on CPU cycles this would save cloud storage systems... Can someone help me quantify the potential savings?
Edit:
They specifically don't list their storage medium on their marketing:
Let's call it "Ring -1"
S3 didn’t used to be strongly consistent, though surprisingly they delivered https://aws.amazon.com/about-aws/whats-new/2020/12/amazon-s3... which I hope they’re proud of.
Some people have been crazy enough to store tables of padded data in the keys of a lot of zero-length objects (which they do charge for) and use ListObjects for paginated prefix queries. It doesn’t much matter whether keys have slashes or commas or what.
In order to maintain high availability they deliberately trade away latency.
So this blog only really applies to local filesystems not object stores like S3.
S3 is similar, in the sense that it has completely different usage than a file system (no hierarchy, direct access, no need for efficient listing etc) so I'm pretty sure they use something similar.
On the other hand, the other facility in the lab had consulted Logica on storage of similar data, who viewed x-y (or multiple dependent variables) data as tables suitable for storage in their early RDB, Rapport. That wasn't actually used in production, and the storage format for the table model was unfortunately usually mangled by data acquisition systems writing files.
In reality there are other operations you wanna do on your files. Concurrent access, updates, deletes are probably not as great. Backing up a database with terrabytes of binaries inside is also non-trivial. Backing up terrabytes of files is a lot simpler.
So absolutely it is a network problem which means custom fiber?
I don't know enough about custom fiber to know whether that will help stretch past being network-bottlenecked — most NICs max out at 10 gigabits/second, but I've heard of faster ones. Eventually you might be able to make yourself disk-limited... Either way, backing up one file is probably easier than backing up a zillion files scattered around the filesystem.
[1]: https://www.cyberciti.biz/faq/linux-unix-apple-osx-bsd-rsync...
Was this the tool & thread you mean? https://news.ycombinator.com/item?id=34303497?
https://learn.microsoft.com/en-us/sql/relational-databases/b...
Not sure how much use it gets though, I never hear anyone talk about it, including Microsoft folks.
To maybe expand on it a bit: it's not known how the author created, configured and mounted the filesystem they were using, neither make nor version of filesystem is given. The author of the test doesn't even know that any I/O test worth its salt needs to "warm up" the system for the results to be reliable. Instead, they run the test multiple times and average them.
To give examples of things that may influence the speed of such tests:
* Is the filesystem mounted with atime option? If it is, it will generate more I/O as open() now generates writes beside reads. This is just one example of mount options affecting the test.
* How is the filesystem configured to store directory info, sometimes it's possible to embed this info in the inode, other times it's a linked-list-like structure that will potentially generate more I/O requests. This is, again, but a single example of a category of factors that have to be controlled for.
* Whether filesystem supports journaling, snapshots, deduplication, compression... whether it's parallel, how big it is / how fragmented it is... how much memory is available for caching, how is system I/O merger configured? And the list of questions not covered by the authors of the test goes on.
The explanation the authors themselves came up with for the results they see is this: the way they designed filesystem tests, they call open() and close() a lot, and they don't do that in their database tests. But, open() and close() don't have a fixed "price". Their performance will depend on many factors listed above and the options given to open().
From what I can tell, the I/O is performed in a blocking way, in a single thread, which is the worst way to perform I/O if you want good performance. So, in this test, both SQLite and filesystem suck, and, if you wanted to make them go faster, then you definitely could. Especially, you could improve the filesystem case.
I work in test automation. In the general storage area. I worked for some years with a distributed filesystem (think something like Lustre, but modern and fast), then worked with something like DRBD, but, again more modern. Not sure how fast (never ran any benchmarks on DRBD). I had to deal with filesystems like Ceph's filesystem, BeeGFS...
Anyways. When I worked on the DRBD analog, let's call the product "R", one of my tasks was to figure out how well would a database work on top of R. Well, "database" is a very broad term. I figured I'd concentrate on using couple of well-known relational databases. PostgreSQL turned out to have the most to offer in terms of insight into its performance. Next, I'd have to find a suitable configuration for the database. And that's where things got really complicated. To spare you the gory details: the more memory you can spare for the benefit of your database server -- the better, the more you can relax the requirement of synchronizing with persistent storage (i.e. fsync and friends) -- even better.
In the end of the day, I had to abandon this kind of testing because, essentially, if given enough memory, or enough replicas (allows not to care about destaging to persistent storage) the bigger numbers I could produce, which made the question "how well does R compare to a plain block device?" irrelevant.
---
Fast-forward to this article. It gives out a vibe of "filesystems are not efficient, if you re-arrange something in the logic of doing I/O you can gain more performance!". And that really reads like "10 things doctors don't want you to know!" advertorial. It's similar to the misguided idea I often encounter with people who aren't system programmers that "mmap is faster".
Now, understanding why "mmap is faster" is nonsense will help understand why benchmarking a database on top of a filesystem and comparing performance doesn't make a lot of sense. So, in order to properly compare the speed, we need to make sure we compare both good and bad paths. What happens when I/O errors occur when memory-mapping files and when using other system API? I invite you to explore this question on your own. Another question you need to ask yourself in this comparison is "how well does this process scale", on a single CPU, single block device, multiple CPUs single block device, single CPU multiple block devices... And on top of this, what if we consider Harvard architecture (to an extend, IBM's mainframes are it, at least there general I/O is separate from the rest of computing). In other words, what if our hardware knows "some tricks"? (Other examples include the kinds of drivers and protocols used to talk to the hardware and whether the storage software, even if running on top of a filesystem will know / be able to take advantage of these, i.e. NVMe allows big degree of parallelism, but will "mmap" be able to utilize that? Especially if multiple files are mapped at the same time?)
And, of course, there is a difference between what's been actually done and what guarantees can be given about the state of data during and after the I/O completes. And these details will greatly depend on the details of the specific filesystem being used. For example, if you wanted to write to the device directly (as in with "ODIRECT | OSYNC"), but the filesystem is something like ZFS or Btrfs (i.e. it needs to checksum your data beside other things), then you might get confused about what you are actually comparing (direct I/O would imply less I/O than is actually necessary to give durability guarantees that may not be given by an alternative storage).
---
So, a better title for the article would have been "It's possible to carve out a use-case, where SQLite works faster than a similar process designed to only use system API to access a filesystem". Which is essentially saying: you, the programmer, don't know how to use a filesystem as well as we do... And, in my experience, most programmers are clueless when it comes to using filesystem or any other kind of storage really. So, that's not surprising, but is still not as sensationalist as the original title.
BTW. There's nothing in SquashFS that makes it "a filesystem in a file". You can use it directly on a device (and it's often used this way), and you can use many other filesystems in a file just as well.
As does:
https://github.com/guardianproject/libsqlfs
https://github.com/narumatt/sqlitefs
(I know nothing about these, just got them from a quick search)