Dat: non-profit, secure, and distributed package manager for data
datproject.org
datproject.org
Dat couldn't be made by nicer people [2]. It's a non-profit led by Max Ogden [3] and the protocol work is led by Mafintosh [4], both of whom are pretty well known in the nodejs community. The project started as a way for academics to work on tabular datasets, but they pretty quickly found out that academics more often want to work on unstructured (or custom-structured) files. So Max recruited Mafintosh and they started working on the p2p protocol, which focused on improving archival and data-sharing flows within labs.
There's a pretty simple CLI you can install from NPM (npm i -g dat). Give it a try. Also, they need help securing grants, so if you have a talent or a connection in that area, definitely get in touch with them. They're doing good work.
1. https://beakerbrowser.com/docs/inside-beaker/
2. https://datproject.org/team
Dat Project started with a focus on increasing access to research & public data. To support the data tools we built a peer-to-peer protocol. People are doing some really cool stuff on top of Dat (such as Beaker Browser), we're really excited about it and want to make sure to support all the neat use cases.
We'll be launching an updated site soon to highlight more of the work around the protocol and what the community is building. Our main use case will still be data management but most of what you see on the current site will shift to a new domain.
Underneath all the transfers in Globus use GridFTP. Using Dat could help distribute bandwidth and speed up transfers. It'll also add version control for free, which (I think) Globus does not have yet.
I use this with encryption for my data folders on my projects.
The core difference is in our approach. We're all open source and a non-profit. We're also really focused on the research data use case where BitTorrent is less easily deployed.
We hope that making an open and easy to use p2p protocol will enable other developers to build applications on top, and something like Resilio could be one.
The direct HTTP support is not fully supported yet. But we're excited for that because it'll allow you to use S3 or other static file servers as peers.
Dat works over any protocol, so it's just a matter of implementing it.
1. Quilt is for profit while Dat is non-profit.
2. Dat has ~20 datasets that are public Quilt has 50+ that are public
3. Dat is on a shared network while Quilt is hosted on a centralized server
4. Both of them offer version control and hosting. Quilt has private hosting for a fee. Dat seems to have only public hosting
5.Quilt is funded by YC. Dat is funded by non profits.
6. Quilt has a Python interface while Dat has one in Javascript
I understand who Quilt is targeting but I'm having trouble understanding who Dat is targeting?
I won't speak authoritatively on behalf of the Dat team, but I believe one of their goals is to make it difficult for public scientific datasets to be lost, and data living on a centralized server is particularly vulnerable to that.
https://beakerbrowser.com/docs/inside-beaker/other-technolog...
I spent a while trying to download recent updates to the Reddit comment corpus [1], which is hosted on BitTorrent. The downloads never seem to finish.
It seems to me that decentralization means that, when a dataset stops being new and exciting, it will disappear. How will Dat counter this?
[1] https://www.reddit.com/r/datasets/comments/65o7py/updated_re...
To counter that, you can take measures to mirror important datasets with a dedicated peer. It requires effort, but it at least makes it much, much harder for example, for a government agency to take down public data without warning.
First, we are working with libraries, universities, or other groups with large amounts of storage/bandwidth. They'd help provide hosting for datasets used inside their institutes or other essential datasets.
Second, we started to work on at-home data hosting with Project Svalbard[1]. This is kind of a SETI@home idea where people could donate server space at home to help backup "unhealthy" data (data that doesn't have many peers).
Finally, for "published" data (such as data on Zenodo or Dataverse), we can use those sites as a permanent HTTP peer. So if no data is available over p2p sites then you can get it directly from the published source.
As others said, decentralization is an approach but not a solution. It gives you the flexibility to centralize or distribute data as necessary without being tied to a specific service. But we still need to solve the problem!
[1] https://medium.com/@maxogden/project-svalbard-a-metadata-vau...
Dat is targeting similar users to Quilt. But we are also looking more broadly at libraries, labs, or other larger academic/gov't organizations managing data. There are a lot of data publishing tools in the sciences such as Zenodo. We'd love for it to be easier to download/publish data to those places. Dat is decentralized, so it really fits well in integrating other data tools.
You can use Dat to replace file transfer software like rsync, so it is a bit more general purpose.
Another difference not mentioned is that Dat really starts at the protocol level while Quilt is more software-focused. Dat protocol is a peer-to-peer protocol for syncing files, modeled off Git and BitTorrent. We built the data management software on top of the protocol.
Edit: I should mention we don't offer any hosting right now, all the data up there is temporarily cached. There is public hosting via Hashbase[1] from the Beaker team. The cool part about Dat being p2p is that it's really easy to switch hosts or use multiple hosts.
Also, is your peer to peer network able to be attacked by nefarious users like a sybil attack? Is there a situation where I could alter or forge data?
The hosting provider is responsible for removing illegal content. Dat itself doesn't track any content. The datproject.org is more of a registry, not a host.
> Also, is your peer to peer network able to be attacked by nefarious users like a sybil attack? Is there a situation where I could alter or forge data?
No, only authorized people can write to each dat key (currently only the owner, but multi-writer is coming soon). All the writes are signed with the writers private key and then verified whenever content is downloaded.
Academics, open data enthusiasts, hackers
It looks like an interesting way to store/share backups and server images - easily scaling bandwidth and availability with the nerd to spin up new instances, or shifting across data centers?
Maybe also as an apt back-end similar to:
I’m not familiar with DebTorrent, but if you’re interested to learn more about the innards of Dat, this post by pfraze is a good place to start:
https://beakerbrowser.com/2017/06/19/cryptographically-secur...
For example, I should be able to give unique URLs (for the same data) to different users and expire one but continue for the other,etc.
Edit: The keys are very short (64 bytes), so they can easily be copy/pasted, tweeted and what have you :)
You generate one key pair for one shared item. I'm talking about multiple key pairs for each shared item, so that I can give access to individual users for the same data and revoke them when necessary.
Fundamentally if you don't have that, then your case is trivially solved by a private bittorrrent tracker.
This is the fundamental difference between things like Quilt and bittorrrent.
http://www.libtorrent.org/dht_store.html
previous HN discussion - https://news.ycombinator.com/item?id=12257065
What you need to mimic dat is a more integrated way to tell other peers that the torrent changed and to check the new one... which is not there yet.
http://www.bittorrent.org/beps/bep_0046.html
>The intention is to allow publishers to serve content that might change over time in a more decentralized fashion. Consumers interested in the publisher's content only need to know their public key + optional salt. For instance, entities like Archive.org could publish their database dumps, and benefit from not having to maintain a central HTTP feed server to notify consumers about updates.
You are technically right that the torrent file is immutable, but basically this lets clients know that the torrent is updated using the DHT data. The outcome is the same.
IPFS seems to me to be a bit over-engineered whereas dat is a lot more simple/low level - something that suits my way of working really well.
Could this be used as a "distributed archive" to store any kind of information? Or does Dat only store some type of data? For example, could I have a "Dat dataset" containing a local Git repo, and every time I update my repo the new files are distributed to whoever is "watching" my dataset?
Under the hood, there are two implementations: hyperdrive and hypercore. Hyperdrive is a filesystem for storing any kinds of files. Hypercore can store any kind of data and is really great for streaming data (hyperdrive is built on top of it).
I was thinking of sharing a Git directory, directly from my computer instead of using any centralized provider such as GitHub/NotABug.
> Hypercore can store any kind of data and is really great for streaming data
is there any web "bridge" available. I'm thinking in particular about streaming data over the P2P network, but also accessible on the web (for example for streaming a video).
I also love simpler things. I'd be curious to understand where is IPFS over-engineered and where is Dat better, instead.
[1] https://news.ycombinator.com/item?id=14771406
EDIT: Annnd...looking up, I see others have already referenced quilt in other comments here.
None of that is true for Dat
How so? As far as I know, git doesn't treat large files any differently than small ones.
As far as I know git isn't good at storing binary data. Git depends on line breaks to be able to diff and make change sets. If you store a binary file in git and make an update to it - even though that update only changed 1 byte, the entire new version of the file is stored again. Dat uses Rabin fingerprinting to intelligently slice binary files into chunks that are less likely to change. That make dat a lot more efficient at storing, versioning, and syncing videos, images, and other large binary files.
Other than that, Git treating large files the same as small ones is part of the problem, and the reason that centralized extensions exist. You wouldn't want "git clone" to clone the complete history of every large file, while that's not a concern at all for small files.
When comparing to Dat, I don't see how; you can't push any Dat repository to Github.
Other than that, Git treating large files the same as small ones is part of the problem, and the reason that centralized extensions exist. You wouldn't want "git clone" to clone the complete history of every large file, while that's not a concern at all for small files.
You don't need centralized extensions for that, though. I use git-annex, which is completely P2P.
Good luck with git-annex.
If the suggested solution is to use git-annex, I will say that git-annex is so poor in usability that the majority response to "you can get the data via git-annex" is "oh well, I'll try some other data then".
Please?!