A notable quote: SQLite does not compete with client/server databases. SQLite competes with fopen().
A notable quote: SQLite does not compete with client/server databases. SQLite competes with fopen().
That is the case, but the advantage of the SQLite approach is that they have already done so, and hence their code is far more likely to be "correct" than any new code you write. The many billions of instances of SQLite already deployed also help with maturity and finding corner case issues.
And as a bonus it also lets you do rollbacks, transactions etc which aren't easy if trying to implement them yourself.
https://en.wikipedia.org/wiki/Files-11
It could work. Just best to have hybrids with different types of files, including those bypassing RDBMS function, so one can select proper reliability vs performance tradeoffs. SQLite might have cross-platform, FS-type API's too that I don't know about. Not sure if it's already there or be an extra development.
http://www.vmssoftware.com/index.html
They're working on it as we speak. :)
The the writes to the `foo.gz` file have to hit the disk before the unlink but the I/O scheduler can reorder these potentially and a badly timed crash could result in data loss. Note that journaling doesn't fix this issue because the transactions are distinct too.
Some days I really do wish we weren't so inured to the particular 50-year-old systems-programming abstractions.
There was a great lightning talk a few years ago that I can't seem to find where the author described an API for storing blobs in a hierarchical namespace. Of course, halfway through, it became clear that it's just the POSIX API: you can "open" handles to objects, "rename" them, remove them, and so on. You'd end a transaction with "fsync()". (Okay, that one's a little more complicated, but I don't think it's as hard as the OP claims, at least for single files. Multiple files are more complicated, but that problem is intrinsically more complicated.)
1. "MVCC": the ability to effectively get a handle on a static copy of the entire filesystem, perform mutations to many objects, and then submit the transaction, at which point it will either commit or rollback depending on whether any of those objects have been modified.
Windows actually has this: https://en.wikipedia.org/wiki/Transactional_NTFS allows for exactly the sort of MVCC I'm talking about—and presents a very different API than the POSIX filesystem one.
2. "Object store": as in, you don't interact with the filesystem by getting writable handles to the file's underlying backing store, where processes can effectively treat a file as persistent shared memory. Instead, you ask for a floating unnamed "buffer" object, fill it up, and then submit it to the filesystem to be atomically stored; or, vice-versa, you ask to retrieve a file, and are passed a buffer handle of the representation of the file at the point in time you asked for it (presumably, for optimization's sake, backed by copy-on-write mmap'ed pages from the latest MVCC-linearized copy of the file.)
We actually have this today as well, but in an implicit and half-assed way. Files below some size threshold, in modern filesystems, don't use any FS extents for backing, but are instead stored as part of their filesystem directory-entry node. This basically makes the filesystem into an object store for those files—but without actually exposing any guarantee to userland that those files will be read/written as atomic operations.
---
To be clearer, there's a problem with this fs-atop-object-store design, as I've stated it so far:
An object store doesn't support every use-case a filesystem does, and a filesystem implemented only in terms of an object-store wouldn't be good for some things because of that. It would be terribly expensive to emulate a block device on top of an object store, for example—you'd have to make each emulated block an object, and mutate blocks by transactionally adding an updated block and removing the old block. This means that it wouldn't make sense to keep mountable read-write disk images on such a filesystem; nor would it make sense to keep database backing stores there.
But both of those things are effectively things that take no advantage of existing on top of a filesystem in the first place—they do their own index-building, their own sparse-allocation, their own journaling, etc. Making the filesystem an object store just makes it obvious that for those use-cases, the filesystem itself is pure overhead.
So, along with the MVCC object-store, the other part of my hypothetical system is a "buffer store": basically like a logical-volume manager, an API that consumes physical block devices and exposes handles to (cheap) logical resizable block devices. The object store could be implemented in terms of one of those logical block devices, but otherwise would completely ignore the existence of it. The filesystem compatibility layer, on the other hand, would allow you to request that a given filesystem object be backed by a persistent buffer (newly-created logical block device) instead of an object; and then whenever you requested that object from the filesystem, you'd get an IO handle to the logical block device, instead of an IO handle to a transactionally-resubmittable copy-on-write mmap of some object.
I think we could have built something that worked well enough, but they grew cowardly after the general failure of the entire longhorn project. Don't try to change the compiler, the language, the os, the office application to use same "in motion" components, plus the database.
It's probably little known outside of Microsoft, but they tried to build many different higher end object filesystems, including JAWS (Jim Alchin Windows Filesystem), and others I can't even remember.
One of my recently-planned hobby projects was to write a FUSE filesystem for OSX that would take all the shell-level glop (LaunchServices UTI-associated verbs, NSFileManager attributes, Spotlight metadata, etc.) and stick it onto the files themselves as plain xattrs, so POSIX tools could effectively be made to work with shell-namespace "Documents" (including treating directory-packages as files, and traversing archive files and container-file-formats as if they were directory-packages.) It'd be quite interesting to combine the two experiments, actually; it'd result in something very much like a POSIX compatibility layer on top of a WinFS-like filesystem.
You don't like files? Files are great.
Running Core Data over SQLite instead of a "real database" gets you backups for free, the iOS file-based security model, lets you delete an app's data when it goes away, network mounted home directories… really, the filesystem abstraction is pretty good still.
So if I need to create a file with a timestamp, you are suggesting creating a SQLite database, a table, and inserting my timestamp there?
It's nonsensical to "fix" a Unix low-level syscall issue by using SQLite.
Most C programmers can't even write that piece of code. Let's see if you can.
It's worthwhile to use libraries which embody the knowledge, trials and errors of years, such as SQLite.
Unless you are talking about a function in the SQLite library that implements the durability calls properly and can be called from other programs.
Stand on the shoulders of the giants who have come before us.
Also, middleware shouldn't be needed in open source because you can simply patch the underlying API. That the middleware is the best solution in this case says a lot about the problem.
I can't believe the level of ignorance in this thread.
But even in an extreme example like storing a single timestamp, there are situations where you want various guarantees that SQLite makes easy.