How Merkle Trees Enable the Decentralized Web
taravancil.com
taravancil.com
"A Merkle-DAG is similar to a Merkle tree in that they both are essentially a tree of hashes. A Merkle tree connects transactions by sequence, but a Merkle-DAG connects transactions by hashes. In a Merkle-DAG, addresses are represented by a Merkle hash. This spider web of Merkle hashes links data addresses together by a Merkle graph. The directed acyclic graph (DAG) is used to model information. In our case, modeling what address has stored specific data. "
(https://www.cio.com/article/3174193/healthcare/from-medical-...)
That would be cool. You just have a partial order between children and parents, like in git.
Can you summarizd how the actual algorithm is supposed to work, ie its main ideas?
I am surprised to find that using sha hashes in a tree or DAG will leak information to people who shouldn't have it. Is this a serious flaw and how do these guys ultimately solve the problem? I read the paper but the algorithm seems a bit hard to understand and follow.
To do that, you need to generate a collision. You enter a random content hash in the document, then keep modifying the document and hashing it untill it's hash matches the one you included. It's not impossible, but depending on the hash function, it may take you a very long while to find it.
Some disadvantage of OP's method:
- hash code is difficult to read for human. Urls under some hierarchy share some common patterns, it also bear some meanings. All hash code will look same and have nothing to hint on the content.
- You have to copy it, almost impossible to type it, or compare two visually similar string.
- You may end up with some link shortening service for hash code, but you can use link shortening to solve the portable file host problem already.
The Merkle Trees can solve some problems, but I don't think portable urls are the right one.
One place I could see it being valuable is with an online archiving type service like Archive.org where the content doesn't change except when a new snapshot of that content is recorded displaying any changes made to that content.
Why? Well a lot of what I think is touted on the IPFS homepage, but to put it in my own words, I've become dismayed over how mutable the web is. It seems to entirely benefit those who seek to lie and disorient. Yet, if something from an "honest person" leaks onto the internet (nude photos, credit card, email, etc) it's nearly impossible to remove it.
This has less to do with the technology, and more to do with human nature. Facts are easy to mutate and spread misinformation about. Pages can be edited, blocked, DDOSed, etc. Yet often leaks of information that small actors, i.e. I upload my secret key, is basically permanently stolen on the web. So I feel like the mutable web is all cost to the public, with no benefit.
I feel/hope that IPFS and IPNS can allow for software to present a normal Reddit-like experience, where users never deal with weird hashes and etc, but underlying it all is an immutable and entirely audit-able paper trail. Information is key in this day and age, and if we can have an immutable web with no UX loss, I think it's a boon.
As it stands, people have identified trends among bad actors on the web - such as politicians botting Reddit to sway opinion, but the trail goes fuzzy quite quick when the entirety of the bots content can be deleted, mutated, etc.
Anyway, just sharing my thoughts / rambles.
Anything I leak is basically permanent on the current web. If that's the case already, why would I want a mutable system in place where people can edit what they said? Alter what was posted, alter votes, take ddos content, etc.
Blahah’s comment above may interest you:
This also avoids the danger of hash collisions, although it is vulnerable to third-party hosts taking down the content (although again, there's not much cost to that. Just move the content elsewhere and change your redirect link).
With this method you do still have some server responsible for hosting the content rather than distributing that load, but that's actually a saving when looking at the storage load across the entire ecosystem.
It's a cool article and I really like the idea of content addressing instead of host addressing. I just feel like it's too divergent from the web as most people use it today (and I'm mildly concerned about malicious hash shenanigans), whereas the above method can be used and understood right now by many web users.
That will be google voice for url. The OP's method is use people's name as phone number.
Another problem with OP's method is that it's difficult to read for human, and very difficult to verify by eye. There is also no common parts for files organized in same hierarchy.
(Also, git keeps the full data, not only diffs)
Merkle trees can work just fine for dynamic content. You just need to use a pubkey pointer.
1. https://beakerbrowser.com/2017/06/19/cryptographically-secur...
IPFS does this by storing (mutable link, contentID) pairs in the DHT, which can be updated if you know the private key for the mutable link.
The chance of that happening is roughly equal to the chance of a collision randomly occurring somewhere in a few quadrillion SHA256 hashes.
https://crypto.stackexchange.com/questions/52261/birthday-at...
http://www.npr.org/sections/krulwich/2012/09/17/161096233/wh...
https://www.cnet.com/news/the-milky-way-is-flush-with-habita...
Hash functions tend to be broken gradually and publicly, and we migrate to new ones as they start to look shaky. It's theoretically possible for someone to privately break a function that everyone else thinks is secure, but it would be an extremely impressive achievement since lots of full-time cryptographers work on breaking these things and publish every little bit of progress they make.
Also, it's highly likely that flaws will be found with the hash algorithm long before 12 million years are up.
Obviously we wouldn't use the same hash algorithm and setup for 12 million years, but the sheer absurdity of that length of time and that pace of content production shows this method will last, at least until flaws are found in SHA256.
I think even 0.01% would be enough to harm usability.
It isn't.
> but why?
"The web" is a collection of independent networks providing a means to traverse hyperlinked text documents across any network that is addressable.
Not only are servers not a problem, host-based addressing is what makes the web work at all, as numerical addressing by itself would have killed the web's growth long ago, and users would have no reasonable way to address content.
In this sense, the web works exactly how any application platform does. You can't get content out of old apps, or apps that no longer run, or on systems that no longer run. This isn't data siloing, this is just legacy applications.
With web apps the data is siloed, but it doesn't have to be, and didn't used to be. The Internet Archive is proof. Getting the data out isn't that difficult, IF it 's not hiding behind a web app. The difficult part is to convince people to stop writing applications which are siloed and do prevent easy access to content that the web used to provide.
Peer to peer networks are not a solution to incompatible legacy applications. It's like a vehicle which is immobile when it runs out of gas. Instead of building the vehicle so it can still be used when it lacks power, they're changing the way the roads work. It's ridiculous.