Past year has had a lot of improvements
- Protocol moved to a hole-punching DHT for peer discovery (hyperswarm).
- Protocol now scales # of writes and # of files much better. We were able to put a Wikipedia export, which is millions of files in 2 flat dirs, into a single "drive" and get good read performance. This performance bump came from an indexing structure that's built into every log entry (hypertrie).
- Protocol now supports "mounting" which is a way to basically symlink drives onto each other. Good composition tool, esp useful for mounting deps in a /vendor directory.
- Browser now has a builtin editor that splits the screen for live editing. Feels similar to hackmd.
- Browser added a bash-like terminal for working with the protocol's filespace. It's glued to the current page so you can drive around the web using `cd`.
- Browser added a basic identity system. Every user has a profile site that's created automatically and maintains an address book of other users.
- We built out application tooling a fair amount. It's fairly easy to build multi-user apps now, where previously it was a bit of rocket surgery.
Some of the year was spent prototyping ideas and throwing them away as well. A bit inefficient, but helped us learn.
Does the newest version of dat handle large files well (10gb)? Does it handle tons of files nested in a few directories well? Are there any issues I should know about there?
What is the command line support like for multi-writer?
Do you have any metrics for how much Dat is currently being used?
Thanks!
Large files work fine but currently any change to a file rewrites it in its entirety. That will mean history will be large until the GC kicks in, and any file that's modified has to be redownloaded in its entirety.
The team spent a fair amount of time looking at a solution to partial file-updates which works like inodes. They ultimately decided it was too difficult to pull off for now.
> Does it handle tons of files nested in a few directories well?
Yep, no issues there
> What is the command line support like for multi-writer?
We're still deciding on how to handle multi-writer. It's a priority for us after the upcoming stable release.
> Do you have any metrics for how much Dat is currently being used?
Nothing concrete atm. If I had to guess, it'd be no more than 1k.
I'd be curious if anyone wanted to try it ;) https://github.com/mafintosh/hyperdrive/blob/master/index.js...
> What is the command line support like for multi-writer?
There is an experimental multiwriter CLI using hyperdrive and kappa-db (github.com/kappa-db)
Speaking of which, the gun.js network hosts the internet-archive (archive.org) [1] which is likely to be bigger than the English Wikipedia (?). Cloudflare runs an IPFS gateway. Are there organizations of similar size dat is looking to onboard or has onboard-ed?
---
A few questions:
1a. Re: hole-punching: Do you support clients on LTE (or, networks behind CG-NAT)? Is hole-punching just for the DHT connections (which seems to operate over UDP4) or for Feed (hypercore) connections, too (which default to TCP?).
1b. Are there plans to use WebRTC (like WebTorrent) instead of TCP or UDP sockets?
1c. In case of TURN relays, how long do you intend to support traffic through your servers? It might get cost-prohibitive at some point. Are (will) TURN relays (be) deployed worldwide to combat latencies?
2. Are hyperswarm and hyperdrive different example usages for hypercore or are they complementary? If so, how are those projects related, as in, is hyperdrive built on top of hyperswarm?
3. The project is spread across disparate GitHub accounts-- Some top-level packages in mafintosh's account and some in hyperswarm's, for instance. Is this intentional?
4. I saw references to merkle-trees, flat-trees, kademlia, HAMT, bitfields... is there a documentation on why these data structures are used? I can guess kademlia is for the DHT and merkle-trees for some form of entropy management. Are such implementation details documented (looking for something similar to gun.js [2] or redis [3] docs)?
5a. Can you please point out a few major differences to gun.js and ProtocolLab's IPFS?
5b. Conversely, what are some use-cases for which dat really outshines other such protocols?
[0] https://news.ycombinator.com/item?id=15818856
Main project website: https://siasky.net
If you want to download skynet files using the command line, you can run 'siad' in portal mode as a daemon, and then use 'siac' to query the daemon. 'siac skynet download [skylink] [destination]' is the command to perform the download.
Doing it this way means you gain direct decentralized access to Skynet instead of needing to funnel through a portal.
In typical p2p networks, you are expected to be both in roughly equal proportions, which means your network gets polluted by a lot of low-quality users who are just trying to do their part.
Sia has the consumers pay the producers, and has a marketplace mechanism that selects the highest performing producers to perform the jobs and receive the revenue.
This allows the network as a whole to be much, much more efficient.
That said, I think there's definitely a space for paid internet. I was just surprised because siasky homepage says "Build a Free Internet" - perhaps it should say "Build a pay per view internet" instead?
I would argue that the current model is more friendly to attention grabbing content (clickbait, etc) because advertising has the limitation that all views are worth the same value. In a pay-as-you-go Internet, users can whitelist high quality content sources as being okay to charge 10x or 100x what you'd typical accept to view a webpage. This would incentive content creators to build a brand and reputation that makes users comfortable putting them on the 'high quality' list, so that their content can see a massive revenue multiple relative to the number of eyeballs.
There are certainly a lot of those, but they are not nearly the whole internet. A lot of interesting sites would not want to be on the "pay as you go model" at all!
For example, we are on the news.ycombinator.com in the thread discussing datprotocol.com, and you are pointing me to siasky.net. Right now, the top 5 HN pages are adecentralizedworld.com, handsonscala.com, arxiv.org, a16z.net, and torproject.com. None of those websites make money from webpage ads. None of them are likely to move to pay-as-you-go internet -- because they care about their audience and not website revenue.
I suspect that Sia's decentralized pay-as-you-go world will be much worse (productivity-wise and information-wise) than the current internet -- all the interesting technical/science blogs and docs would be missing; while buzzfeed clones will be plentiful. There might be occasional high-quality journalism website, but most of those are getting paywalls anyway, so will they justify all new protocol?
Also, the internet is already something that users have to pay for. Users pay their ISP every month for access. We envision a world where this utility payment extends to cover the content creators in addition to the infrastructure providers.
Disclosure: I work at Cliqz.
Does that mean that Cliqz will start indexing dat sites ? The way I see it, the biggest issue of decentralized anything is always discoverability. If Cliqz is interested in it, it could have a leverage in making dat more widespread.
Dat even helps you with the indexing, because it's the sites themselves that tell you when they are updated and when they are updated, it's like an indexer's dream come true. And seeing how dat operates, Cliqz could even be a peer in the swarm of the sites it indexes to give them more chance to be reached (at the cost of Cliqz having a view on who visits what, which seems contrary to your DNA if I understand correctly)