Hyperdrive v10 – a peer-to-peer filesystem
blog.hypercore-protocol.org
blog.hypercore-protocol.org
To make SQLite decentralized (like Hyperdrive) you can put in a torrent. Index it using full-text-search https://sqlite.org/fts5.html for instance. Then let the users seed it.
Users can use sqltorrent Virtual File System (https://github.com/bittorrent/sqltorrent) to query the db without downloading the entire torrent - essentially it knows to download only the pieces of the torrent to satisfy the query. This is similar techniques behind Hyperdrive I believe just again, using standard tools and tech that exists and highly optimized: https://www.sqlite.org/vfs.html
Every time a new version of the SQLite db is published (say by wikipedia), the peers can change to the new torrent and reuse the pieces they already have - since SQLite is indexed in an optimal way to reduce file changes (and hence piece changes) when the data is updated.
I talk a bit about it here: https://medium.com/@lmatteis/torrentnet-bd4f6dab15e4
Again not against redoing things better, but why not use existing proven tech for certain parts of the tool?
So now to update the torrent file you need a mechanism for having a mutable document you can update in a distributed but signed way. Or you could make an append only feed of sequential torrent urls... oh wait.
My point is: Hyperdrive's scope is sufficiently different from your proposed solution that yes, you could probably rely on existing tools (and I have much love for bittorrent based solutions!) but it starts feeling like shoehorning the problem into a solution that doesn't quite fit.
That draft status is of little practical relevance, though, if nothing changed for years, and no one voiced well-founded critic on the technical details.
I do agree though that Hyperdrive is different from what the bittorrent ecosystem has to offer. I too like not reinventing the wheel where that's not necessary, as you recommend there. I'll leave you the list of BEPs for further reading, in case you're interested: https://www.bittorrent.org/beps/bep_0000.html
This is also why sqlite is a good choice because it's highly optimized to do the least amount of changes to its "pieces" when an update occurs.
If you're implementing this behavior, trying to manage all kinds of different queries, building a querying engine on top of that, optimizing for efficiency and reliability, you're effectively rewriting a database. Sure you can do it, but why not take advantage of battle-tested off-the-shelf stuff for things like "databases" (sqlite) and/or "distributing data" (torrent)?
I'm going to read your blog post now, thanks a lot for the new info.
With dat-sdk, users just need to go to a webpage. You really just need WebRTC without torrents.
I got rid of multiwriter by just having a dat archive for each user, and the users sharing their dat addresses with each other. They write to their own. When that happens, events emit and users listening write to theirs.
If enough users stay on, listening to each other's address, I only need a web client.
Also, if I have offline support, like Workbox Background Sync, I don't even need internet and information transfers device to device with just an offline PWA. At least that's my goal.
> As of June 2015, the dump of all pages with complete edit history in XML format at enwiki dump progress on 20150602 is about 100 GB compressed using 7-Zip, and 10 TB uncompressed.
From: https://en.wikipedia.org/wiki/Wikipedia:Size_of_Wikipedia#Si...
It's a bit of a nitpick either way because you're right, Wikipedia may not be the best example because 10TB is still relatively small.
I've been looking around IPFS, dat, hyperdrive etc and it seems like dat is the most natural setting for this but sqltorrent is new to me.
Can you please mention your thoughts on:
- Discoverability of content in Hyperswarm (DHT Search/"Superpeers"/???)
- What happens to the old DEP proposals? (There is a critical feature that I need for my service that's still an open DEP proposal! https://github.com/datprotocol/DEPs/issues/61)
- What uses cases do you have in mind for this service?
Thanks!
Not sure if I'm answering your first point, but Hyperswarm is baked into the Hyperdrive daemon, so daemon users should have their drives swarmed/available automatically. There are a few CLI commands to toggle this behavior too, in case you don't want to add your drive's key to the DHT.
The Hypercore Protocol org's creating a similar proposal repo called HYP [0] (we couldn't resist the name), scoped tightly to the core protocol. We're still solidifying the proposals plan, but yours would add lots of value, so we don't want to lose track of it.
Since your proposal is about peer identifiers, you might like that Hyperswarm uses the Noise protocol [0] to handshake each connection, and the Noise key can be used as a stable peer ID.
As for use-cases, check out Beaker (launched today too). Paul's made a whole bunch of example applications that take advantage of Hyperdrive features, like drive mounts.
I'm personally really interested in using Beaker to make personal document indexers + search (kinda like an amped-up Dropbox), but that's for another post!
I tried to install it and got some c compile errors. Probably my nodejs version is too old, but you want to have the absolute minimum of friction for users to install it.
Beaker comes as an appimage for linux, which worked flawlessly at the first try. Maybe do that for the hyperdrive daemon as well?
Last time Dat was on HN I tried to follow the "simple chat application" tutorial[0], but got stuck at the stage where 2 instances were supposed to automatically discover each other because they only intermittently managed to actually discover each other.
Will this new version of Hyperdrive improve this? Or is it something completely different?
There are some diagrams describing this on the Hypercore Protocol [0] site, and the repo's can be found here.
We also talk a bit more about this on the hypercore website
How POSIX-like is it?
From a filesystem I would expect some more rigorousness than just stating this without any extra information, especially as implementing an actual, to-spec, fully distrubuted POSIX file system is known to be a difficult problem and is generally solved by not wanting a full POSIX-compatible implementation instead (ie. no locking, append-only files, etc.).
While reading Hyperdrive/Hyperswarm docs I actually managed to find everything I wanted to know without much effort (most of all it seems like setting up this on private network should be doable)
Hyperdrives also have a built in capability system where you have to know the public key of the drive to download it from a peer, so if you only share the key with yourself no one else can access it.
Finally using the modules you can build almost any kind of networking you’d imagine but that part requires more work on your side obviously.
You don't have to worry about a flood or tornado taking out all the photos of first birthdays, or Grandma when she was younger than you are now, and you also don't broadcast family dynamics, whereabouts (people have been robbed when social media made it clear they were not at home) or pictures of minors.
You, your uncle who Moved to the City and your cousin by marriage who wants to be a game designer all set up a file server and share the photos (your cousin is gonna throttle traffic while he's playing CoD or Dota of course, which it turns out he does all the time but at least he's an extra backup copy).
I don't know how you keep that one relative from uploading funny things they found on the internet that keep trying to install spyware or back doors though, but I suppose you'd have that problem now.
I think we have all the tech needed already (WebRTC, DHT, NAT hole punching, decades of p2p, encryption, onion security, StorJ/Filecoin) etc but what is lacking is dead simple UX and wide support of operating systems - Windows, Mac, Android, iOS, Linux (Raspberry, Synology, cheap VPS backup)
Still looking for the perfect Dropbox-like experience but without the cloud sync piece. If only syncthing had decent mobile apps...
I'm still waiting for a 'Drobo' like device without the proprietary physical layout. A light or dial goes into orange territory, you head to Best Buy or Amazon and buy the biggest drive that doesn't give you sticker shock, you push a button, out pops your worst drive and in goes the new one. Some lights flicker for a while and then go green.
I thought ZFS would have given us almost everything but the hardware ten years ago, but it turned out they oversold a few of the features back then, and then Larry happened.
Custom hardware is too expensive for small run consumer hardware, and Apple might have gotten into that space but never did. I wonder how many PCIe lanes you could shoehorn onto a Pi clone...
A freenas box does this as well, but won't have the pretty drive light indicator if it's a home-built box, but then you're not limited to proprietary hardware.
For the physical device at least, Debian and 2 Btrfs data drives in RAID 1 certainly isn't turnkey but seems quite accessible at this point.
I'm looking into integrating with DAT or Hyperdrive-like solutions to help make storage backups less fiddly for users.
For now, I recommend my beta users use SyncThing or Resilio Sync to get their photos and videos off their phones and on to their home NAS or computer.
The only major drawback seems to be that you have to host the physical hardware yourself due to lack of solid end-to-end encryption for most platforms. Sandstorm might have it (I'm not clear on what's client- and what's server- side there), Seafile has end-to-end encryption that doesn't protect metadata (https://forum.seafile.com/t/how-strong-is-the-encryption/627...), and NextCloud appears to have a long-running beta of end-to-end encryption on a per-folder basis (https://nextcloud.com/endtoend/). Apparently Cryptomator exists (https://cryptomator.org/) although I've never tried it myself.
If you want to decentralize your shared files unfortunately last I checked SyncThing didn't yet support end-to-end encryption. (I don't think any of the other ones I mentioned can be used in a decentralized manner but things move quickly so I'm not sure.)
Alternatively, if you were thinking more chat and messaging there's self hostable federated services such as PixelFed, PeerTube, Mastodon, and Matrix. Or was there some other usecase you had in mind?
2 questions:
- What is the difference between dat and hyperdrive ?
- I see hyperdrive manages hyper:// urls -- what about dat:// urls ? Will they be still managed ?
- "Hyperdrive" was formerly the internal data-structure name of Dat archives. With this release, the team decided to rename the protocol from Dat to Hypercore Protocol. Subsequently, "Dat Archives" are now "Hyperdrives." The Dat community will post some updates about this soon.
- The hyperdrive-daemon does not support dat:// URLs. I don't know what the future of dat:// URLs will be but Beaker is phasing them out with a converter tool.
Beaker uses Hyperdrive, it's basically the source of its novel features. The Beaker team works on Hyperdrive (the Hypercore Protocol) but we maintain a separate org at https://github.com/hypercore-protocol
Take a look at the "mounts" section of the blog post where we describe a group pattern you can set up. You can create a group directory called "team-drive", then mount each user's drive into the group.
If you're using FUSE, this will feel similar to Dropbox, but with one directory per-user.
We're starting with these kinds of simple mounts, and brainstorming ways to extend them soon.
For a full “union” mount experience we still have some research to do but we are working on it.
The mount setup is really good though and fully p2p
Now, onto the content, you touch on de-duplication. I am quite concerned with the cost associated to updating a large file. Is something like rolling hashes investigated, to chunk files independently of their size? I guess it is, given you seem to be working hard on de-duplication.
But then, that kind of trick best works on uncompressed data, which is inefficient for transmission. Is data compressed before transmission? Whole chunks, or whole files at a time? Ahead of time? Interactively based on what the peer needs?
The trade-offs are many, and complex to investigate. I'm wondering if this could be used as an OS image, like OSTree does?
And lastly, I did not get if multiple peers having the same private key identity could modify the structure simultaneously. What would happen?
Also, nodejs gave me a kneejerk reaction that may be unwarranted, but that's quite a huge dependency to pull in for something that wants to be a ubiquitous building block. Does it have a C API? Also, twitter, discord, github (node to a lesser extent)... it seems somewhat ironic to build a ultimate decentralized filesystem while relying on these hypercentralized offerings, and I am afraid it could turn some contributors off.
But dat/hyperdrive is a well-documented protocol, and there are several ports to other languages.
I am very interested in the rust port, but I am not sure in what state it is: https://datrs.yoshuawuyts.com/
I have seen some tweets about progress being made on this. Does anybody know more?
The wire protocol works now: https://github.com/Frando/hypercore-protocol-rs and the community is active in #datrs on freenode
1. Consensus. If I submit conflicting updates and sign both, how does the swarm resolve what is the latest state?
2. Migrating the swarm. If all the machines in a swarm get corrupted, can I migrate to a totally new swarm?
- Hyperdrive includes mass-publishing as a usecase, so it uses public-key URLs and a bandwidth-sharing mechanism among its active peers (like BitTorrent)
- Hyperdrive is built on a general-purpose protocol called Hypercore which is a signed append-only log. These logs can be used for other datastructures. Some examples [1] [2]
[1] Kappa-core, a general-purpose db built on the logs https://github.com/kappa-db/kappa-core
[2] Cabal, a chat network https://cabal.chat/
node_modules/hyperdrive-daemon/node_modules/hyperdrive-daemon-client/bin/commands/create.js:8
static usage = 'create [path]'
^
SyntaxError: Unexpected token =On a slightly related note, is anyone interested in having a discussion about how to layer on top of all these "drive" systems a `HDFS` like drive that would have n block replication across different sources, and trying to interpret any given source (dropbox, hyperdrive, google drive, etc..) blocks to make sense of what is being stored there would render the person confused?
:)
As for 'network questionable' and mobile, here are a few of the other projects building on Hypercore that have made those a priority [0] [1] [2].
The Hyperdrive daemon as it's currently built wouldn't fare too well in a bandwidth and/or battery constrained environment (wasn't designed for that), but a mobile solution is definitely on our radar.
How would hyperdrive deal with the following scenario: you got a large dataset such as wikipedia. Lots of people have browsed it, but most of them only have a tiny fraction.
How do yo know which peers to connect to to get a particular bit (offset?) you are interested in? The DHT only tells you which nodes participate in the hypercore, not what they have in detail, right?
At the moment we don't do anything special in regards to discovery, but as we scale that's something we want to investigate. Since everything is running on append-only logs we can group the data into sections quite easily so there is some easy wins we can do there with announcing to the dht that you have data in a specific region.
I had looked at https://datprotocol.github.io/how-dat-works/ , but I don't remember anything about a gossip protocol or a peer building a "view of the world". Is that new?
We are working on expanding this scheme so peers can help discover peers that have the section you are looking for.
Due to the compressed bitfields these section are quite large. In most cases using a few kilobytes you can share WANT/HAVE for millions of blocks
Being able to identify a piece of content by an integer instead of a hash makes things more efficient compared to content-addressed storage a la IPFS.
Hyperdrives builds a p2p filesystem on top of Hypercore for a single writer. Using mounts you can mount other peoples drives so merge conflicts don't happen since there is no overlapping writes.
We are working on a union mount approach as well for overlapping drives (we talk a bit about this in the post)
Do you agree that for the collaborative data structures side of things (like the chat app) users of the hypercore-protocol will likely run into clock trust problems?
PS I'm a big fan of your work/repos.
But I would not call it finished, as of today, apparently only one person can make changes to the filesystem. That does limit the use cases.
"In v10, we don't go all the way to a general multi-writer solution; solving multi-writer scalably, without incurring major performance penalties or confusing UX, remains a research question for us. "
.. so, give them some support, so they can solve this.
[1] hyper://1bc1faf01a22270fb5698a60e63ef7a596ad976457e6d9914a8fd56d87281917/
The page you are trying to view cannot be shown because the authenticity of the received data could not be verified.
I don't know much about this space, but always been interested in some sort of web socket/WebRTC p2p fs.