DwarFS – Deduplicating Warp-Speed Advanced Read-Only File System
github.com
github.com
DwarFS: A fast high compression read-only file system - https://news.ycombinator.com/item?id=32216275 - July 2022 (64 comments)
DwarFS: A fast high compression read-only file system - https://news.ycombinator.com/item?id=25246050 - Nov 2020 (111 comments)
Feature request: Add a "library" option to give mkdwarfs a list of files that should be loaded into the dedup mechanism first, but not stored, allowing the image to be even smaller if the contents of the file can be retrieved from that library instead. Bonus points if you can specify a dwarfs image as a library and have it sensibly use the files contained in it.
Then you have the basis for a deduplicating incremental backup system. Currently, I have a system I wrote that will take a single file and a list of library files and produce a compressed deduplicated file that can re-create that single file using the library, which is great if you use tar to create that single file, but a little unwieldy when coming to decompress and restore everything. The bonus of making it a proper mountable filesystem instead is that then it's a proper mountable filesystem and retrieving single files is a doddle.
My use case is that I have students, and I have given them coursework, which involves them logging in to a Linux machine and hacking away. I want to store regular snapshots of their work so that I can keep a backup for their sake but also so I can see a progression of development to try to work out if they are cheating (yes, I have had to deal with this), but I don't want to store 100 copies of the same fairly large files.
It has a great system to snapshot files but only store data if it's changed. I use it in an environment where I can't use something like zfs to snapshot data because I don't have the ability to make decisions about what filesystem we're using. It's been amazing, love it so much!
> my main use case and major motivation was that I had several hundred different versions of Perl
okay... The question is if you want to trust a github filesystem or just spend the $dough to deal with it.
I’m thinking a hybrid approach could make sense here. Historical data is thrown into a big archive dwarFS while new data of e.g. the current day is kept in a simple normal folder. Every now and then this would be merged together.
Skimming through the docs I can’t tell if it will be possible to recreate a new dwarFS by taking an existing one and adding just a few more files, or if I would need to create a completely new one which also means I need to temporarily have enough space for twice the archive size
To me it seems such an fs should be immutable-first rather than read-only, i.e. let you create (by copying from another fs) files which just can't be changed (can be moved though) ever after. And such would be what I actually need as I have a lot of big and relatively redundant files which aren't meant to change (e.g. video files, picture originals, distro and backup disk images etc.).
$ mkdwarfs -i /path/dir -o image.dwarfs
[1]: https://github.com/mhx/dwarfs/blob/main/doc/mkdwarfs.md
As then, you can dedup per block, even with a sliding window. Nicest is that you can do it fully asynchronously, so real time appends are not slowed down.
Which are many.
Some with deduplicarion built in.
...Or, in the author's case, when you have hundreds of directories that have only small differences compared to each other.
As other posters in the bigger thread mentioned, this is also very useful for big arcade game collections.
The same way we [used to] use specialized software tools to assemble an ISO 9660 image with our data, and then burn that image to a [single-session] write-once optical disk like CD-R.
would maybe a database dump be smaller?
>As of May 2015, the current version of the English Wikipedia article / template / redirect text was about 51 GB uncompressed in XML format.
Compressed data at the same time was 11.5 GB. And that's data from 9 years ago, and just English Wikipedia.
For comparison, I collect leaked password dumps and they (combined, after deduplication) go into hundreds of GBs too. And that's for just username:password lines, not even text.
This kind of feature is really important if you to encourage people to seed even after their download is completed.
Since ipfs is content-addressable storage, it's readonly. If we make it cow, encourage people to seed is easy.
And then they announced they will willingly and proactively delete any hash that any legislative agency tells them to, and it was dead in the next minute.
> "You need to mount the fuse drive with the option "allow_root"." https://serverfault.com/a/1085669
Did I misunderstand?
Thanks
It's been a while since I did that comparison, so the results could be significantly different now.
[1] https://www.amazon.com/Ontrack-Data-International-99-00030-0...
Would it be possible to take the core design changes here and apply them to squashfs, and maybe propose a next major version of the squashfs internal format to make all these things possible?
https://news.ycombinator.com/item?id=39417503
SeaweedFS supports WebDAV. https://github.com/seaweedfs/seaweedfs/wiki/WebDAV
I'm not able to find if both/restic supports mounting backups as WebDAV, but in theory there's nothing stopping you.
It's 100% user space (expose a rest service) and supported by a bunch of file-browsers with a bit of a network aware component to it as well.
However, since it's a FUSE only file system, it's difficult to see how it would be used on embedded system firmware, so it could perhaps see use as a distribution mechanism. Similar to tar or zip files, but possibly with (much) better performance for random access, should you need only smaller portion of the whole archive.
The author indicates need for keeping multiple similar copies of sets of unchanging files on their computer, and made this to reduce the space needed for them, while retaining the access through the file system. So that is also a use case.
Neat instruction.
"Clustering of files by similarity using a similarity hash function" does anyone have an intuition of how similarity hashes work?
But the self-congratulatory tone from description is… unusual, to say the least.
I actually only went through a significant portion of the readme (through the CromFS part) because of your comment, and i just don't see it. I see a person who wrote actually useful software that is multi-platform and gives the positives and negatives of the software they wrote compared to alternatives available today.
In every test DwarFS compared favorably, and on tests where one aspect was marginal, the DwarFS code was better in other regards: power at the wall, extract/read times, etc.
How would it be better presented by a solo developer?
I can almost hear “and if you call now, you get this amazing towel for free!”
But judging by the comments here, it seems I’m more sensitive to this tone than most.
> HAProxy offers a fairly complete set of load balancing features, most of which are unfortunately not available in a number of other load balancing products
It seems like a minor nit that could be "fixed", especially since the "it gets better" is in reference to "not only smaller, but also much faster" and it's not an insignificant performance increase, it's 100 times faster. Basically everyone (mostly) uses squashfs, and this absolutely trounces it - according to the author.
anyhow i hope my reply wasn't too extra
...as if... now, let me spit out all this purple democratic kool aid