Dat – Distributed Dataset Synchronization and Versioning
github.com
github.com
Some interesting properties:
1. It uses a BitTorrent-style of swarm, but primarily to sync signed append-only logs, which are in fact flattened Merkle Trees (similar to Certificate Transparency). The Dat archives are addressed by public keys. The tree is used to enforce the append-only constraint by making it easy to detect if the history has been changed by the author.
2. The "Secret Sharing" feature. The public key of a Dat archive is hashed before querying or announcing on the discovery network, and then the traffic is encrypted using the public key as a symmetric key. This has the effect of hiding the content from the network, and thus making the public key of a Dat a "read capability": you have to know the key to access its files.
There's a reference implementation in JS available at https://github.com/datproject/dat-node, and a fair number of tools being built around it.
Flask is already used ( http://flask.pocoo.org/ )
Maybe Retort? Funnel? Or maybe a proper name? Berzelius? Erlenmeyer would be difficult for people to pronounce.
I maintain ConceptNet [1], a multilingual knowledge graph. I do everything I can to make its published results reproducible. The biggest hurdle for people reproducing it has always been getting the data -- building it requires about 100 GB of raw data or 15 GB of computed data that can be imported into PostgreSQL.
I once tried git-annex. It turned out not to be a good choice -- its tools were flaky, its usage patterns confusing, it leaves a permanent record of your mistakes in configuring data sources, and it was very hard to convince to use ordinary HTTP downloads instead of trying to get read-write access to S3 (which wouldn't work for anyone but me). Now I have weird branches and remotes in my repositories, and weird data in my S3 buckets, that I can't get rid of in case someone tries to use git-annex in a way I told them would work.
After that I just went with distributing the data with plain HTTP downloads from S3. I wish I could do better than this. The only semblance of versioning is putting the date in the URL, and also people in Asia tell me that the build fails because their downloads from us-east-1 get interrupted. Oh, and if I ever stop paying for S3, everything will break.
If I tried making data reproducible with Dat, would it be safe to promise people that they could use Dat to get the data? Even if in the future I don't like Dat anymore?
For instance, do I have to commit to hosting the data somewhere? If not, who does? Does it disappear when people lose interest, like BitTorrent?
Dat's really similar to BitTorrent when it comes to availability; it doesn't do anything automatically to guarantee it. If you choose to use Dat, you'll need to ensure a peer exists, though public peer services will be available soon.
Dat's still young and you'll probably have to endure some hiccups, but if you do want to give it a try, PM me and I can help you get started.
It sounds like Dat is a ways off from being something I could use as an authoritative source of data, but I could include it as one way to get the ConceptNet data. If it succeeds, that could save on the S3 bill and maybe even distribute the data across continents better. (Heck, I'm sure a lot of the downloads are scripts I'm running, and I could be file-sharing the files to myself.)
And I guess a Dat URL shouldn't really be an authoritative entry point referred to in a paper, as it would just point to one version of the data with no context. Maybe Dat plus Zenodo could do the trick eventually.
Your points are spot on and things we've been thinking a lot about. I also wouldn't feel comfortable putting a Dat link in a paper yet, but that is an eventual goal because of the persistence properties of dat compared to http urls.
> Maybe Dat plus Zenodo could do the trick eventually.
Yes, exactly!! Dat supports http publishing right now, so you can run `dat sync --http` on a server and it'll live publish files from the source to an http site.
We are working on also supporting http downloads. The idea is you publish to Zenodo, including the SLEEP metadata, dat can then clone the files over http (including content verification) or via the peer network, i.e. `dat clone zenodo.org/record/439922`.
We are super excited for the http downloading because it'll allow dat to store files on any data repository with a good api and a http file backend. We've been talking with the Dataverse folks on how to accomplish this there and have an eye on others such as Zenodo.
Awesome, congratulations! :):)
For distributed/p2p software these are just amazing times to be alive :)
Just today I was using the multilingual Conceptnet-numberbatch word vectors[1], which would not be possible without your work.
To your point though - you can use Amazon S3 as a seed for Bitorrent downloads, which might help some and reduce what you pay. See [2]
[1] https://github.com/commonsense/conceptnet-numberbatch
[2] http://docs.aws.amazon.com/AmazonS3/latest/dev/S3Torrent.htm...
I have some background in question answering over knowledge graphs, though, so I'm familiar with their strengths.
In my company Luminoso's work, it's important in building domain-specific models that can be used for topic detection, search, and classification. Beyond that, I use it for mostly the basic demos -- word similarity, text similarity, analogies, et cetera.
I believe based on its performance there that it should be a pure upgrade to the kind of applications that use word2vec, but I'd like to know what particular applications it's being used in besides my own.
100GB is way too large for the project I'm working on at the moment (dbhub.io), as even a bunch of people downloading something that large would nuke our sponsorship budget since we're just starting out (still pre-launch).
However, if we gain traction and become cash positive, data sets this size would be good to cater to. :)
Switching to PostgreSQL sped things up, at the cost of requiring a separate database process, dealing with psql's weird access control, and adding an inconvenient step of loading the data using COPY commands.
TBH, I'm not sure how a site can work without mutability.
All DHTs suffer from this issue. It's just particularly likely that IPNS's use of a DHT will lead to attacks. Making it not susceptible to this would require a redesign of IPNS's protocol.
The two features I list above, the append-only histories and secret-sharing, are unique to Dat. And, for us, the URL spec was a big deal.
IPFS still has the long-term goal of path addressing (NURI), we just hadn't yet completely figured out what the upgrade path should look like. The discussion linked above is turning into a spec and into actions, i.e. IPFS will be trying to get as much as possible of Electron's protocol API [2] into WebExtensions, [3] and make us of that in the browser addons. [4]
[1] https://github.com/ipfs/specs/pull/152#issuecomment-28462886...
[2] https://electron.atom.io/docs/api/protocol/
Here's why I think you're shooting yourself in the foot with the NURIs, though.
1) You're focusing on syntax. The "there is no domain" premise works just as fine with your stage two of `ipfs://{hash}`, so dismantling domains doesnt really justify the NURI change. The path syntax, of `/ipfs/{hash}`, contains functionally the same information.
2) If you still need IPNS, then you still need a concept of domains, so the "there is no domain" premise isn't really accurate.
There's the concept of nestability in NURIs that is supposed to increase protocol composition, but I think you're overgeneralizing, at the cost of breaking backwards compatibility. When we looked at using IPFS in Beaker, the NURI was a major problem for us, because we're limited to what Electron/Chrome provides. It's not just API surface either: there's a lot of code in Chromium that makes assumptions around the standard URL syntax. Are you really so sure that nestable references are worth the headache? Because you're gambling the entire IPFS project on it.
One of the reasons I like the IPFS project so much is that they produce a lot of good side products (multiformats, libp2p) that can help anyone build an IPFS-like system.
So I wouldn't see it that bleak. Even if that choice of NURIs is a fatal flaw as you claim, it is still a very surface-level problem that could be fixed in a fork of it.
- You can publish new versions to a URL until you somehow forget the private key, and then it's fixed forever, so long as people hang onto copies.
- There's nothing to prevent people from passing around a URL with a version in it. So, although it looks like the author has some control, this is an illusion; publishing is irrevocable and anything published could go viral. (This is generally true of making copies, but it's the opposite of Snapchat.)
- Suppose someone chooses to publish a private key? Is it a world-writable URL? Hmm.
At the moment that would result in each leaked-key user maintaining a different history, with different peers only downloading the updates from the leak-author they happen to receive data from first. But in the future what will happen, once we get to writing the software for it, is the corruption event will be detected and recorded by all possible peers, freezing the dat from receiving future updates.
It would be good to figure out key rotation for Dat URL providers since this probably has to be built into the protocol.
Any thoughts on integrating with keybase? I like keybase's model where you have device-specific keys. But this would probably make moving a Dat URL provider to a different machine trickier.
This all assumes that well-known Dat URL's become an important thing to preserve (they are published in papers, etc) even though they are very user-unfriendly, even more than IP addresses.
A naming system on top of them would make key rotation a non-issue (rotate Dat URL's instead) and you could completely replace or remove the history, sort like a git rebase. But that loses other nice properties of the system?
I suppose irrevocability is something we deal with in git repos all the time. Although you can do a rebase locally, once a commit is accepted by a popular project, they're unlikely to let you remove it from history. The review process makes it unlikely that any really embarrassing mistake would be accepted, so this seems ok in practice.