Freqfs: In-memory filesystem cache for Rust
docs.rs
docs.rs
And more on topic: tokio-uring is really fast [1], and I'm really loving tokio in general.
[1] https://gist.github.com/munro/14219f9a671484a8fe820eb35d26bb...
For example, most file systems today are journaling. Which is exactly how most databases handle atomic, consistent, and durability in ACID.
About the only thing it's missing is automatic document locking (though most file systems support explicit locks).
That said, there are often some pretty hard limits on the number of objects in a table (directory). Depending on the file system you can be looking at anywhere from 10k to 1 billion files per directory.
There are also some unfortunate storage characteristics. Most file systems have a minimum file size of around 4kb, mostly to optimize for disk access. DBs often pack things together much more tightly.
But hey, if you can spin using the FS as a DB... Do it. Particular for a read heavy application, the FS is nearly perfect for such operations.
Astronaut with gun: always has been.
https://en.wikipedia.org/wiki/WinFS
Or more like Beos BFS with its extended attributes, indexing and querying?
https://en.wikipedia.org/wiki/Be_File_System
Also I think a lot of the old mainframe filesystems had the concept of records and indexes built in since they were primarily used for business operations.
The record based approach had many properties we know from modern databases. It was a first class citizen on the mainframe and IBM was its champion.
In my opinion hierarchical filesystems won as everyday data storage because of their simplicity and not despite it. I think the idea of a file being just a series of bytes and leaving the interpretation to the application is ingenious. That doesn't mean there is no room for standardized OS-level database-like storage. In fact I'd love to see that.
It's the difference between synchronous indexing that's baked into the system (as in file system metadata structures and database indexes, which update at the same time your data is changed) vs. fragile add-ons that index asynchronously (which in general I find tend to be too slow to update, missing results, and prone to breaking).
perhaps filesystems should be extensible in a way that supports indexing intelligently.
(i know you mentioned an aversion to addons)
the file as an opaque box for applications to store a real data structure is poisonously anti-file. it's totally what files are, what we think of them, but imo, systems like 9p, or linux's procfs or sysfs are The True Way for files: small discrete pieces of data which are part of a system of directories tlthat express a larger compilated hierarchical system of data.
Files won, but only the stupidest wrongest version. Easy to copy and manage but utterly useless on their own, unscriptable, pointless eithout their complex applications there to use them.
I dont think db's/records are that interesting either. i think we just need to really try files. Fine grained files. As opposed to these big ole blobs the OS cant really interact with.
I would prefer more pluggable interfaces personally.
(hi Ryan, long time no see!)
There was some research being done on the concept of a db as a filesystem: https://youtu.be/wN6IwNriwHc
And I think ReiserFS was also working towards this but got abandoned for obvious reasons.
Try it here: https://github.com/blobcity/db
PS: I am the chief architect of the DB, and the project is no longer being actively maintained by us. But if you make a contribution, we will oblige to review and merge a PR.
Bottom line, nothing you do can make your database faster than the filesystem. So why not make a database that just uses the filesystem to the fullest, than creating a filesystem on top of a filesystem. BlobCity DB does not create a secondary filesystem. It dumps all data directly to the filesystem, thereby giving peak filesystem performance. This is scientifically really the best it gets from a performance standpoint. Not necessarily the most efficient in data storage / data-compression standpoint.
This means, we gain speed, while compromising on data-compression. We produce a larger storage footprint, but are insanely fast. Storage is cheap, compute isn't. So that should be okay I suppose.
Previous HN discussion: https://news.ycombinator.com/item?id=20394088
Why not let the OS take care of this?
Maybe your caching strat of you OS isn't best for your use case. Also, you may use a network file system, or several types of FS, and want your cache warm up to be tuned up and consistent.
On the other hand the OS does know about memory pressure from IO and from heap memory for the whole system. This crate will only know about cache pressure within a single process.
> Also, you may use a network file system
Which can also be set to do aggressive caching, at the expense of consistency.
> and want your cache warm up to be tuned up and consistent.
the description doesn't say that it's doing cache warmup any more eagerly as regular reads would
Sometimes you need very explicit control over when things are read from cache and when they aren't. This can be hard with network file systems. Especially when you have two different use cases on the same filesystem, which isn't that odd, even within a single application.
The major advantage of freqfs over just letting the OS handle file caching is that with freqfs you can read and mutate the data that your file represents purely in memory. For example, if you implement a BTree node as a struct, you can just borrow the struct mutably and update it, and it will only be synchronized with the filesystem in the event that it's evicted (or you explicitly call `sync`). This avoids a lot of (de)serialization overhead and defensive coding against an out-of-memory error.
Again, I will update the documentation to clarify.
Love it when a program could simply work, but chooses to fail because it doesn't like my life choices.
This library might have other uses that I am not aware of.
https://github.com/serprex/openEtG/blob/master/src/rs/server...
edit: seems it can. Nice