DwarFS: A fast high compression read-only file system
github.com
github.com
I've looked at DwarFS before for this same use case but the fact that it's read only makes it more difficult to handle since I'd have to have an uncompressed version sitting out elsewhere. Though I've now got the hardware to actually put that all into a CI/CD that generates the image. I might actually work on that once I get my turing pi 2 boards since I want to port this whole setup to ARM as well as x86_64. I might also put it on my RISC-V board too but it's too slow so I think it won't be as useful.
EDIT: fixed the tbd size, it was taking a while to calculate.
So I'd imagine making something like that but with DwarFS below would be quite easy although it'd require you to set it up by hand. Still, once done it'll likely be a rock-solid setup for a long time.
Credit and thanks to coldblues for alerting me about this!
N.B. there is also EroFS https://www.kernel.org/doc/html/latest/filesystems/erofs.htm...
handle = vfs_fopen("/my/file1.txt")
pointer_to_the_file_bytes = vfs_map(handle, <start offset>)
which would be as fast as possible. compression, encryption aren't needed.For “as fast as possible” you'll need to experiment and benchmark with your workload. Which filesystem is optimal may depend on how you are laying out the data and where the latency/throughput sensitivities are in your use case and the given filesystems.
> My prime concern is having an api such as...
Having a different API other than it looking like a filesystem would make cross-platform more of a concern, as you then have a data access library not a general filesystem. It will likely to be necessary for best performance though: any filesystem is going to have significant overheads (orders of magnitude) compared to being able to map chunks of the data directly into your process' address space.
If abandoning a generic filesystem, perhaps something like sqlite with an in-memory table/db (https://www.sqlite.org/inmemorydb.html)? Again like the ramdisk option just load up the content from permanent storage on first use.
A normal unprivileged app cannot create RAM disks easily on any OS as far as I know ; also it wouldn't really work on e.g. WASM
> as you then have a data access library not a general filesystem.
That's fine for me - although I don't see a particular difference between either, a filesystem is just a system to access files, whatever that means
The distinction I'd make us that a filesystem provides a common generic API that practically all processes on the OS understand and share access to. Pretty much always implemented out-of-process (in the kernel or another userland process via kernel stubs/hooks like FUSE).
A data access library is usually much more specific to a particular data set or set of applications, and likely doesn't follow the filesystem abstraction (at least not in the same way).
There's also a vfs module for it that imitates a filesystem on top of a single SQLite DB.
It's also what "powers" a wide range of portable filesystem in a single file tools such as DOCX and ODT and quite a few other modern Office and Office-adjacent file formats.
Or is it that you create a compressed 'file' that you can mount as a file system? Like a zip file kinda I guess?
The source data can be a raw block device or more likely a local file. It doesn’t matter as either way the kernel is just reading blocks of bytes.
Not really tbh. I haven't used a CD in my adult life. I remember "burning a CD" was a thing and that's about it.
I was born in 1991 and I haven't used one since I was in 9th grade and made a girl a 'mix tape'.
Somebody else can chime in with the exact mechanism by which this one is written, but common solutions include being writable sometimes or having a program to build the filesystem from known data. That might be filesystem-as-a-file, filesystem on a separate partition, or what have you.
What resources would people here suggest for learning about file systems? I see a lot of new file systems like zfs, btrfs, etc. I looked up for resources but couldn't find anything substantial.
I want to learn how they work so that I can appreciate projects like this and compare them.
I looked into the Build Your Own X repo but didn't find anything. I found a book called Practical File System Design: The BeOS file system but it's apparently dated, and I'm not sure I want too much of a deep dive.
> The Super Simple Storage Service (S4) is a new innovation in cloud storage. Our advanced write-only storage provides the highest security, lowest cost, and simplest management available.
Does that mean that I can store all those 100s of photos taken at slightly different angles, in an efficient way?
Most compressed filesystems will fail to dedup stuff because for various performance related reasons the compression window is usually quite small maybe up to 128kb and the files are often not sorted in the archive such that small files would manage to compress anyway in that window.
EDIT: Apparently squashfs does include file de-dupe however DwarFS is fixing what I basically said above. It sorts files in order of similarity so that you can compress accross file boundaries as well as just removing complete duplicates. That's pretty cool.
"Clustering of files by similarity using a similarity hash function. This makes it easier to exploit the redundancy across file boundaries."
This kindof thing is more difficult to do on a writable filesystem though not impossible.
Shame that it can't be used much commercially in many cases due to GPLv3 license. This could improve so many embedded systems currently using SquashFS.
(the dwarfs dev would have to work it out amongst all who contributed to his repo)
What I'm saying is that due to the license fewer companies want to use and thus we're getting a bit worse products as a consequence.
Of course I completely agree the developer should get licensing income (heck, that's in my self interest as well!), but what's going to really happen is that the cheapo companies just make do with less and use SquashFS or whatever instead.
It's a tragedy of commons there's so little will to donate for or crowdsource open source projects and that they're taken for granted.
Which is not made better by your sentiment? I mean GPL at least forces you to either try negotiating a different license with the author (for money ideally), or suck it and use something inferior. Using a permissive license that allows multi-billion dollar companies to use your code for free certainly doesn't help changing the mind set of "open source is other people working for me for free".
Sounds great in theory, but even for a popular open source library developer this happens so rarely that it rarely pays the bills regardless of license.
"Using a permissive license that allows multi-billion dollar companies"
Most embedded devices by far are not developed by multi-billion dollar companies. Their legal departments are also actively steering away from GPL licenses.
"try negotiating a different license"
In my experience, by far the most companies don't want to negotiate anything. If there's a product with a set price, yes, then it might be purchased. They typically also want some kind of product support with it.
"use your code for free certainly doesn't help changing the mind set of "open source is other people working for me for free"."
How to solve this? Ideally the library developer would get paid AND the consumers get more value for their money. This would encourage more developers to write useful libraries that provide great value in the big picture, but are uneconomical or otherwise too much trouble to deal with on an individual company or product level.
As it is, everyone seems to lose.
If you're including this existing filesystem you shouldn't have any relevant patents to worry about, and the clause about letting the user write their own firmware shouldn't be an issue for 99% of products.
Is there anything else I missed in the differences between v2 and v3?
Use some vendor blob -> problem. Have some NDA for registers used in some hardware -> problem. Hardware does not have user firmware flashing -> problem. Legal compliance requirements for hardware bearing your name -> problem.
That's the same as GPLv2. If you don't mix the filesystem into your proprietary code, you don't have to reveal any of those things.
> Hardware does not have user firmware flashing -> problem.
GPLv3 only requires you to let the user have the same flashing ability you have. If it can't be flashed, you don't have to do anything.
> Legal compliance requirements for hardware bearing your name -> problem.
That's a 1% of products situation.