BTFS: BitTorrent Filesystem
github.com
github.com
Many moons ago I created a Linux distribution for a bank. It was based on Ubuntu NetBoot with minimal packages for their branch desktop. As the branches were serverless, the distro was self-seeding. You could walk into a building with one of them and use it to build hundreds of clones in a pretty short time. All you needed was wake-on-lan and PXE configured on the switches. The new clones could also act as seeds. Under the hood it served a custom Ubuntu repo on nginx and ran tftp/inetd and wackamole (which used libspread, neither have been maintained for years). Once a machine got built, it pulled a torrent off the "server" and added it to transmission. Once that was completed the machine could also act as a seed, so it would start up wackamole, inetd, nginx, tracker etc. At first you seed 10 machines reliably, but once they were all up, you could wake machines in greater numbers. Across hundreds of bank branches I deployed the image onto 8000 machines in a few weeks (most of the delays due to change control and staged rollout plan). Actually the hardest part was getting the first seed downloaded to the branches via their old Linux build, and using one of them to act as a seed machine. That was done in 350+ branches, over inadequate network connections (some were 256kbps ISDN)
I actually have never used it in order to know if AWS puts their own trackers in the resulting .torrent or what
Wouldn't creating new torrents for each update to a dataset cause clients to retransfer data that hasn't changed?
https://blog.libtorrent.org/2020/09/bittorrent-v2/
Especially merkle hash trees which enable :
- per-file hash trees - directory structure
I tried doing something like this in the past using ipfs, but I wasn't very successful and didn't find any duplicates.
Related, I use and recommend https://github.com/cyanreg/cyanrip on modern UNIXes.
https://github.com/thomas-mc-work/most-possible-unattended-r...
Finding a good CD drive to rip them is the first step.
https://flemmingss.com/importing-data-from-discogs-and-other...
IME Discogs had the track data most often.
And obviously rip to flac
The main advice I can give you is to use ripping software that integrates with AccurateRip (XLD, EAC, etc) and use a widely supported lossless format (like FLAC).
Also — I can’t remember all the details, but there’s a way to store a CUE file, along with some metadata alongside your rip such that you can recreate an exact copy of the original physical media.
At least for now, I’ve moved on to streaming services, but I’m happy to know that I have a large library of music that I ripped myself to fall back to using instead, should I ever choose to.
https://github.com/automatic-ripping-machine/automatic-rippi...
I never had it fully working because the last time I tried, I was too focused on using VMs or Docker and not just dedicating a small, older computer to it, but I think about it often and may finally just take the time to set up a station to properly rip all the Columbia House CDs I bought when I was a teen and held on to.
Watch a show do some other work and when the toast pops out a new one in.
Ripping DVDs with HandBrake was almost as easy, but it wouldn’t eject the disc afterwards (though it could have supported running a script at the end, I don’t recall).
Now, I'll start another round with DBPowerAmp's ripper on macOS, then I'll see which tool brings the better metadata.
I used to use magnetico and wanted to make something that would use crawled info hashes to fetch the metadata and retrieve the file listing, then search a folder for any matching files. You'd probably want to pre-hash everything in the folder and cache the hashes.
I hope bitmagnet gets that ability, it would be super cool
I’ve written tools to inspect content (say in an ISO file system), and those will hash to the same value (so different sector data but the same resulting file system). Audio converted to CDDA (16-bit PCM) will hash as well.
If audio is transcoded into anything else, there’s no way it would hash the same.
At my last job I did something similar for build artifacts. You need the same compiler, same version, same settings, the ability to look inside the final artifact and avoid all the variable information (e.g. time). That requires a bit of domain specific information to get right.
(but then the .torrent file itself has to be stored on a storage that resists bit flipping)
As someone with no storage expertise I'm curious, does anyone know the likelyhood of an error resulting in a bit flip rather than an unreadable sector? Memory bit flips during I/O are another thing but I'd expect a modern HDD/SSD to return an error if it isn't sure about what it's reading.
https://documents.westerndigital.com/content/dam/doc-library...
What I understand by bit flip is a corruption that gets past that check (ie the "flips balance themselves" and produce a valid ECC) and returns bad data to the OS without producing any errors. Only a few filesystems that make their own checksums (like ZFS) would catch this failure mode.
It's one reason I still use ZFS despite the downsides, so I wonder if I'm being too cautious about something that essentially can't happen.
I plan to write documentation of the IPFS process including the PXE router config later at https://github.com/majbacka-labs/nixos.fi -- we might also run a small public build server for peoples Flake configs, who are interested in trying out this process.
You're doing some cool things here.
Remember in the late 90's booting server off a CD-ROM was the thing.
But you shouldn't need to: you should be able to do the same thing with a docker graph driver, so there is no registry - even daemon should perceive the local registry as "already available", even though in reality it's going to just download the parts it needs as it overlay mounts the image layers.
Which would actually potentially save a ton of bandwidth, since the stuff in an image is usually quite different to the stuff any given application needs (i.e. I usually base off Ubuntu, but if I'm only throwing a Go binary in there plus wanting debugging tools maybe available, then in most executions the actual image pulled to the local disk would be very small).
I created Spegel to fill the gap but focus on the P2P registry component without the overhead of running a stateful application. https://github.com/spegel-org/spegel
"Kraken was initially built with a BitTorrent driver, however, we ended up implementing our P2P driver based on BitTorrent protocol to allow for tighter integration with storage solutions and more control over performance optimizations.
Kraken's problem space is slightly different than what BitTorrent was designed for. Kraken's goal is to reduce global max download time and communication overhead in a stable environment, while BitTorrent was designed for an unpredictable and adversarial environment, so it needs to preserve more copies of scarce data and defend against malicious or bad behaving peers.
Despite the differences, we re-examine Kraken's protocol from time to time, and if it's feasible, we hope to make it compatible with BitTorrent again."
Fully distributed OS's/Virtual Machines/LLM's/Neural Networks
If LLM's are token predictors for language, what happens when you do token prediction for computation across a distributed network? Then run a NN on the cache and clustering itself? Lots of potential use cases.
btfs https://archive.org/download/BigBuckBunny_124/BigBuckBunny_1... mountpoint
btplay https://archive.org/download/BigBuckBunny_124/BigBuckBunny_124_archive.torrentBTFS – mount any .torrent file or magnet link as directory - https://news.ycombinator.com/item?id=23576063 - June 2020 (121 comments)
BitTorrent file system - https://news.ycombinator.com/item?id=10826154 - Jan 2016 (33 comments)
I do think IPFS is awesome, but is going to take some major advances in at least 3 areas before it becomes something usable day-to-day:
1. not running a local node proxy (I hear that Brave has some built-in WebTorrent support, so maybe that's the path, but since I don't use Brave I can't say whether they are "WebTorrent in name only" or what
2. related to that, the swarm/peer resolution latency suffers in the same way that "web3 crypto tomfoolery" does, and that latency makes "browsing" feel like the old 14.4k modem days
3. IPFS is absolutely fantastic for infrequently changing but super popular content, e.g. wikipedia, game releases, MDN content, etc, but is a super PITA to replace "tip" or "main" (if one thinks of browsing a git repo) with the "updated" version since (to the best of my knowledge) the only way folks have to resolve that newest CID is IPNS and DNS is never, ever going to be a "well, that's a good mechanism and surely doesn't contribute to one of the N things any outage always involves"
I'm aware that I have spent an inordinate amount of words talking about a filesystem other than the one you submitted, but unlike BTFS, which I would never install, I think that those who click on this and are interested in the idea of BTFS may enjoy reading further into IPFS, but should bear in mind my opinion of its current shortcomings
In v2 this is solved and it is possible to easily know the hash of each file in the torrent, so you can search for it in other torrents
Then I don't see much of an advantage over just vanilla bittorrent - if you realistically need a full local copy to even start working anyway.
It lets you create a block device (/dev/nbd0) backed by a torrent. Blocks are fetched in the background, or on demand (by prioritizing blocks according to recent read requests).
In practice it works - you can even boot a VM from it - but it's quite slow unless you have lots of seeds. There's a danger, particularly with VMs, that you can hit time outs waiting for a block to be read, unless you adjust some guest kernel parameters.
There are some bootable examples in that page if you want to try it.
I had no idea what I was doing, most of the hard work IS done by the torrent-stream node package
Feed her some hungry reggae, she'll love you twice
The only usecase I see for this is as an alternative to a more traditional bittorrent client.
Pretty minor to do, but, it's a big speed increase.
A few projects that tackled some version of this...
Nextcloud, Owncloud, and generally just NAS can be "dropbox but self hosted" but its centralized.
IPFS, Perkeep, Iroh, hypercore (npm), focused on content-addressed information, making cataloging and sharing easy, but fail to really handle coordination of which node the data goes on.
Syncthing, Garage, and of course BitTorrent, and a few others can coordinate but they all duplicate everything everywhere.
"Bazil.org", and (dead link) "infinit.sh" both sought to coordinate distribution and somehow catalog the data, but they both seem to have died without achieving their goals. I used infinit.sh when it was alive ~2016 but it was too slow to use for anything.
I'd love for something like this to exist, but I think its an impossible mission.